Llama 4 Scout
Meta's open-weight multimodal model with a standout 10M-token context window and single-H100 viability in Int4.
Pricing
InputOpen weights
Cached inputn/a
OutputSelf-hosted cost profile
Meta distributes weights; serving cost depends on your runtime.
Cost bandFree/open
Free tierOpen weights; hosted pricing depends on partner.
Limits
Context window10M checked Jul 19, 2026
Max outputSee runtime
Knowledge cutoffSee launch materials
DeploymentSelf-hosted, Hugging Face, llama.com, partner clouds
OpenAI-compatible APIYes
ModalitiesText, Image input
Where it fits
Best forExtreme-context experiments, open deployment, and teams building around the Llama ecosystem.
Business fitAttractive when open deployment and context scale matter more than turnkey hosted polish.
Capabilities
Open weightsMultimodalMoESingle-H100 Int4 viability
Benchmark scores
Composite76 / 100
Coding74 / 100 (LiveCodeBench)
Reasoning73 / 100 (MMLU-Pro)
Long context88 / 100 (RULER 10M)
Vision73 / 100 (MMMU)
Instruction following76 / 100 (IFEval)
Editorial aggregate of published benchmark results as of Apr 1, 2025. These are not independent measurements run by this site.
Change history
- Apr 18, 2025New modelLlama 4 Scout released 10M context in a single-H100-viable package. Changes the scale ceiling for open deployments.
Data
Last checkedJul 19, 2026
Cross-checked againstopenrouter.ai/meta-llama/llama-4-scout
Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.