Nemotron 3 Ultra
Open frontier-reasoning and orchestration model from Nvidia: a hybrid Transformer-Mamba mixture-of-experts with 55B active of 550B total parameters and a 1M-token context window.
Pricing
Input$0.60 / 1M checked Jul 19, 2026
Output$3.60 / 1M checked Jul 19, 2026
Hosted-inference pricing as listed; open weights available for self-hosting.
Cost bandBudget
Free tierOpen weights published; hosted API pricing shown.
Limits
Context window1M checked Jul 19, 2026
DeploymentSelf-hosted, NVIDIA NIM, partner inference
OpenAI-compatible APIYes
ModalitiesText
Where it fits
Best forSelf-hosted frontier reasoning and agent orchestration on Nvidia hardware.
Business fitOpen weights with hosted pricing under $4 per 1M output tokens.
Capabilities
ReasoningFunction calling
Change history
- Jun 4, 2026New modelNemotron 3 Ultra added to the index Open-weight 550B MoE (55B active); hosted at $0.60 in / $3.60 out per 1M, 1M context. OpenRouter listing
Data
Last checkedJul 19, 2026
Cross-checked againstopenrouter.ai/nvidia/nemotron-3-ultra-550b-a55b
Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.