Llama 3.1 8B

Meta · Llama 3 · open · released Jul 23, 2024 · meta-llama/Meta-Llama-3.1-8B-Instruct

Llama 3 family:Llama 3.1 405B·Llama 3.1 8B·Llama 3.3 70B

Meta's efficient open model — runs on a single consumer GPU and punches well above its weight class.

Pricing

InputOpen weights / ~$0.05 / 1M hosted
OutputOpen weights / ~$0.05 / 1M hosted
Cost bandFree/open
Free tierOpen weights (Meta Llama license).

Limits

Context window128k checked Jul 19, 2026
Max outputRuntime-dependent
Knowledge cutoffDec 2023
DeploymentSelf-hosted, Ollama, Groq, Together, Fireworks, Hugging Face
OpenAI-compatible APIYes
ModalitiesText

Where it fits

Best forEdge deployments, local inference, and pipelines where cost/hardware is the primary constraint.
Business fitThe benchmark reference for small open models — everything else compares to Llama 3.1 8B.

Capabilities

Open weightsFunction callingEdge inference

Benchmark scores

Composite55 / 100
Coding52 / 100 (HumanEval)
Reasoning54 / 100 (MMLU-Pro)
Long context54 / 100 (RULER 128K)
Instruction following60 / 100 (IFEval)

Editorial aggregate of published benchmark results as of Sep 1, 2024. These are not independent measurements run by this site.

Change history

No recorded changes for this model yet. Pricing and context are checked daily; changes appear here with dates and sources.

Data

Last checkedJul 19, 2026
Cross-checked againstopenrouter.ai/meta-llama/llama-3.1-8b-instruct

Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.

Open in the comparison table JSON API