Llama 3.1 8B
Meta's efficient open model — runs on a single consumer GPU and punches well above its weight class.
Pricing
InputOpen weights / ~$0.05 / 1M hosted
OutputOpen weights / ~$0.05 / 1M hosted
Cost bandFree/open
Free tierOpen weights (Meta Llama license).
Limits
Context window128k checked Jul 19, 2026
Max outputRuntime-dependent
Knowledge cutoffDec 2023
DeploymentSelf-hosted, Ollama, Groq, Together, Fireworks, Hugging Face
OpenAI-compatible APIYes
ModalitiesText
Where it fits
Best forEdge deployments, local inference, and pipelines where cost/hardware is the primary constraint.
Business fitThe benchmark reference for small open models — everything else compares to Llama 3.1 8B.
Capabilities
Open weightsFunction callingEdge inference
Benchmark scores
Composite55 / 100
Coding52 / 100 (HumanEval)
Reasoning54 / 100 (MMLU-Pro)
Long context54 / 100 (RULER 128K)
Instruction following60 / 100 (IFEval)
Editorial aggregate of published benchmark results as of Sep 1, 2024. These are not independent measurements run by this site.
Change history
No recorded changes for this model yet. Pricing and context are checked daily; changes appear here with dates and sources.
Data
Last checkedJul 19, 2026
Cross-checked againstopenrouter.ai/meta-llama/llama-3.1-8b-instruct
Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.