Google DeepMind · launched 21 Jul 2026
Gemini 3.5 Flash-Lite
Google's fastest, most economical 3.5-class model for high-volume agentic tasks, translation and simple data processing.
Read the official announcement- Vendor
- Google DeepMind
- Launch date
- 21 Jul 2026
- Input price
- $0.30 / 1M tokens
- Output price
- $2.50 / 1M tokens
- Cached input
- $0.03 / 1M tokens
- Batch discount
- 50%
- Context window
- 1.05M tokens
- Max output
- 66K tokens
- Input types
- Text, Image, Video, Audio, PDF
- API model ID
- gemini-3.5-flash-lite
Text, image, video and audio inputs share the listed rate. Cache storage and grounding are billed separately.
Official pricing source · checked 8 Oct 2026
Price compared with other models
| Model | Launched | Input / 1M | Output / 1M | Cached / 1M | Blended / 1M | vs Gemini 3.5 Flash-Lite |
|---|---|---|---|---|---|---|
| Claude Haiku 5.5AnthropicListed prices and comparisons apply to prompts up to 100K tokens. Above 100K, input costs $0.50, output $2.50 and cache reads $0.05 per million tokens. | 7 Oct 2026 | $0.10 | $0.50 | $0.01 | $0.20 | 76% cheaper |
| GPT-6 LunaOpenAI | 22 Sept 2026 | $0.10 | $0.50 | $0.01 | $0.20 | 76% cheaper |
| GPT-5.6 LunaOpenAI | 9 Jul 2026 | $0.20 | $1.20 | $0.02 | $0.45 | 47% cheaper |
| Gemini 3.5 Flash-LiteGoogle DeepMindText, image, video and audio inputs share the listed rate. Cache storage and grounding are billed separately. | 21 Jul 2026 | $0.30 | $2.50 | $0.03 | $0.85 | Baseline |
| Gemini 3.8 FlashGoogle DeepMindIntroductory rates through 31 December 2026. From 1 January 2027, input is $1.50, output $7.50 and cached input $0.15 per million tokens. Cache storage and grounding are billed separately. | 2 Sept 2026 | $0.75 | $3.75 | $0.075 | $1.50 | 76% more |
| Claude Haiku 4.5Anthropic | 15 Oct 2025 | $1.00 | $5.00 | $0.10 | $2.00 | 2.4× the price |
| Grok 4.7SpaceXAIBase rates apply up to 200K context tokens; higher-context requests use different rates. Batch API is not supported. The separate fast variant costs twice as much. | 21 Sept 2026 | $2.00 | $6.00 | $0.50 | $3.00 | 3.5× the price |
| Claude Sonnet 5Anthropic | 30 Jun 2026 | $2.00 | $10 | $0.20 | $4.00 | 4.7× the price |
| Claude Sonnet 5.5AnthropicCache reads were reduced from $0.20 to $0.10 per million tokens on 7 October 2026. | 28 Sept 2026 | $2.00 | $10 | $0.10 | $4.00 | 4.7× the price |
| Gemini 4 ArgonGoogle DeepMind · Limited accessAnnounced introductory rates for the upcoming broader launch. Regular rates will be $4 input and $20 output per million tokens; the introductory end date and batch discount are not published. | 30 Sept 2026 | $2.00 | $10 | $0.10 | $4.00 | 4.7× the price |
| GPT-6 SolOpenAIListed rates apply up to 272K input tokens. Above 272K, input costs $4, cached input $0.40 and output $15 per million tokens for the full request. | 22 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 | 4.7× the price |
| GPT-6.1 SolOpenAIListed rates apply up to 272K input tokens. Above 272K, input costs $4, cached input $0.20 and output $15 per million tokens for the full request. | 29 Sept 2026 | $2.00 | $10 | $0.10 | $4.00 | 4.7× the price |
| Gemini 3.1 ProGoogle DeepMind · PreviewRates apply to prompts up to 200K tokens. Above 200K, input is $4, output $18 and cached input $0.40 per million tokens. Cache storage and grounding are billed separately. | 19 Feb 2026 | $2.00 | $12 | $0.20 | $4.50 | 5.3× the price |
| Claude Opus 5.5Anthropic | 22 Sept 2026 | $4.00 | $20 | $0.20 | $8.00 | 9.4× the price |
| GPT-5.6 SolOpenAIPrice before the GPT-6 launch, from OpenAI's announcement. OpenRouter now lists $2 / $10. | 9 Jul 2026 | $4.00 | $20 | – | $8.00 | 9.4× the price |
| Claude Opus 5Anthropic | 24 Jul 2026 | $5.00 | $25 | $0.50 | $10 | 12× the price |
| Claude Fable 5Anthropic | 9 Jun 2026 | $10 | $50 | $1.00 | $20 | 24× the price |
| Claude Fable 5.1Anthropic | 1 Sept 2026 | $10 | $50 | $0.25 | $20 | 24× the price |
| GPT-6 AstraOpenAI | 4 Sept 2026 | $10 | $50 | $1.00 | $20 | 24× the price |
GPT-6.1 Sol and GPT-6 Sol rates were checked on 8 October 2026 against OpenAI model documentation. USD per million tokens, checked 2 Oct 2026 via OpenRouter. Comparison uses the blended price (3 input tokens for every output token). Haiku 5.5 and Sonnet 5.5 prices were updated from Anthropic's 7 October 2026 announcement. Gemini and Grok prices were checked against official Google and xAI documentation on 8 October 2026.
Official benchmark results
Reported by Google DeepMind on 21 Jul 2026 · exact values
Exact published scores · Higher is better. Unreported results are omitted.
| Benchmark | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite | GPT-5.4 mini | Claude Haiku 4.5 |
|---|---|---|---|---|
| SWE-Bench Pro (Public)Diverse agentic coding tasks | 54.2% | 38.3% | 54.4% | 39.5% |
| Terminal-bench 2.1Agentic terminal coding · Terminus-2 harness | 54% | 31% | 59.2% | 44.2% |
| MLE-BenchMachine Learning Engineering | 39.2% | 22% | – | – |
| GDPVal-AA v2Knowledge work · Elo | 1140 Elo | 642 Elo | 1171 Elo | 907 Elo |
| OSWorld-VerifiedAgentic computer use | 74% | 54.3% | 72.1% | 50.7% |
| CharXiv ReasoningInformation synthesis from complex charts · No tools | 74.5% | 73.2% | 80.3% | 61.7% |
| CharXiv ReasoningInformation synthesis from complex charts · With tools | 76.5% | 75.6% | – | – |
| GDM-MRCR v2 (8-needle)Long context performance · 128k (average) | 72.2% | 60.1% | 42.7% | 35.3% |
| GDM-MRCR v2 (8-needle)Long context performance · 1M (pointwise) | 21.3% | 12.3% | – | – |
- Bold marks the best reported score in each row.
- Exact scores from the Google DeepMind model card. Earlier competitor models are preserved as benchmark comparisons only.
- CharXiv reports results with and without tools. Long-context results use separate 128K average and 1M pointwise settings. A dash means no result was published.