Gemini 2.5 Flash
Google's price-performance model for low-latency, high-volume tasks that still need reasoning and tools.
Pricing
Input$0.30 text/image/video, $1 audio
Cached input$0.03 text/image/video, $0.10 audio
Output$2.50
Batch pricing cuts input to $0.15 and output to $1.25.
Cost bandBalanced
Free tierFree tier available with limits.
Limits
Context window1,048,576 checked Jul 19, 2026
Max output65,536
Knowledge cutoffJune 2025
DeploymentGemini Developer API and Google AI Studio
OpenAI-compatible APINo
ModalitiesText, Image, Audio, Video input
Where it fits
Best forHigh-volume assistants, multimodal pipelines, and cost-controlled agents.
Business fitUsually the Google default when budget and throughput matter more than absolute model quality.
Capabilities
ThinkingCachingBatch APICode executionFile searchFunction callingGroundingStructured outputs
Benchmark scores
Composite78 / 100
Coding76 / 100 (LiveCodeBench)
Reasoning79 / 100 (GPQA Diamond)
Long context82 / 100 (RULER 1M)
Vision83 / 100 (MMMU)
Instruction following79 / 100 (IFEval)
Editorial aggregate of published benchmark results as of Sep 1, 2025. These are not independent measurements run by this site.
Change history
No recorded changes for this model yet. Pricing and context are checked daily; changes appear here with dates and sources.
Data
Last checkedJul 19, 2026
Cross-checked againstopenrouter.ai/google/gemini-2.5-flash
Verify with provider documentation before committing spend. Report an error and it gets fixed with a logged correction.