Phi-3.5 Mini
Microsoft's 3.8B model with 128k context — excellent instruction following in a tiny footprint.
Pricing
InputOpen weights
OutputSelf-hosted cost profile
Cost bandFree/open
Free tierOpen weights (MIT).
Limits
Context window128k
Max outputRuntime-dependent
Knowledge cutoffMar 2024
DeploymentHugging Face, Ollama, ONNX, Azure AI Foundry
OpenAI-compatible APIYes
ModalitiesText
Where it fits
Best forExtremely constrained edge deployments, mobile inference, and sub-$1/hr compute scenarios.
Business fitBest-in-class at the 3-4B tier; choose Phi-4 if you can spare 10× the parameters.
Capabilities
Open weights (MIT)Long context for sizeEdge inference
Benchmark scores
Composite58 / 100
Coding56 / 100 (HumanEval)
Reasoning58 / 100 (MMLU-Pro)
Long context56 / 100
Instruction following62 / 100 (IFEval)
Editorial aggregate of published benchmark results as of Oct 1, 2024. These are not independent measurements run by this site.
Change history
No recorded changes for this model yet. Pricing and context are checked daily; changes appear here with dates and sources.
Data
Last checkedJan 10, 2026
Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.