Qwen3 32B
Qwen3's mid-range dense model with hybrid think/no-think modes on a single-GPU footprint.
Pricing
InputOpen weights
OutputSelf-hosted cost profile
Cost bandFree/open
Free tierOpen weights.
Limits
Context window128k checked Jul 19, 2026
Max outputRuntime-dependent
Knowledge cutoffSee release materials
DeploymentSelf-hosted, Hugging Face, Ollama, vLLM
OpenAI-compatible APIYes
ModalitiesText
Where it fits
Best forSelf-hosted reasoning and coding where 235B is too large to serve efficiently.
Business fitBest Qwen3 model for teams that can't serve the flagship 235B MoE.
Capabilities
Thinking and non-thinking modesOpen weightsFunction calling
Benchmark scores
Composite79 / 100
Coding82 / 100 (LiveCodeBench)
Reasoning84 / 100 (GPQA Diamond)
Long context70 / 100
Instruction following78 / 100 (IFEval)
Editorial aggregate of published benchmark results as of Jun 1, 2025. These are not independent measurements run by this site.
Change history
- Jul 19, 2026Context increaseQwen3 32B context window now 128k
32kto 128k OpenRouter directory
Data
Last checkedJul 19, 2026
Cross-checked againstopenrouter.ai/qwen/qwen3-32b
Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.