Google DeepMind · launched 2 Sept 2026
Gemini 3.8 Flash
Google's latest stable Flash model for long-horizon software engineering, autonomous agents and complex enterprise workflows.
Read the official announcement- Vendor
- Google DeepMind
- Launch date
- 2 Sept 2026
- Input price
- $0.75 / 1M tokens
- Output price
- $3.75 / 1M tokens
- Cached input
- $0.075 / 1M tokens
- Batch discount
- 50%
- Context window
- 1.05M tokens
- Max output
- 66K tokens
- Input types
- Text, Image, Video, Audio, PDF
- API model ID
- gemini-3.8-flash
Introductory rates through 31 December 2026. From 1 January 2027, input is $1.50, output $7.50 and cached input $0.15 per million tokens. Cache storage and grounding are billed separately.
Official pricing source · checked 8 Oct 2026
Price compared with other models
| Model | Launched | Input / 1M | Output / 1M | Cached / 1M | Blended / 1M | vs Gemini 3.8 Flash |
|---|---|---|---|---|---|---|
| Claude Haiku 5.5AnthropicListed prices and comparisons apply to prompts up to 100K tokens. Above 100K, input costs $0.50, output $2.50 and cache reads $0.05 per million tokens. | 7 Oct 2026 | $0.10 | $0.50 | $0.01 | $0.20 | 87% cheaper |
| GPT-6 LunaOpenAI | 22 Sept 2026 | $0.10 | $0.50 | $0.01 | $0.20 | 87% cheaper |
| GPT-5.6 LunaOpenAI | 9 Jul 2026 | $0.20 | $1.20 | $0.02 | $0.45 | 70% cheaper |
| Gemini 3.5 Flash-LiteGoogle DeepMindText, image, video and audio inputs share the listed rate. Cache storage and grounding are billed separately. | 21 Jul 2026 | $0.30 | $2.50 | $0.03 | $0.85 | 43% cheaper |
| Gemini 3.8 FlashGoogle DeepMindIntroductory rates through 31 December 2026. From 1 January 2027, input is $1.50, output $7.50 and cached input $0.15 per million tokens. Cache storage and grounding are billed separately. | 2 Sept 2026 | $0.75 | $3.75 | $0.075 | $1.50 | Baseline |
| Claude Haiku 4.5Anthropic | 15 Oct 2025 | $1.00 | $5.00 | $0.10 | $2.00 | 33% more |
| Grok 4.7SpaceXAIBase rates apply up to 200K context tokens; higher-context requests use different rates. Batch API is not supported. The separate fast variant costs twice as much. | 21 Sept 2026 | $2.00 | $6.00 | $0.50 | $3.00 | 2× the price |
| Claude Sonnet 5Anthropic | 30 Jun 2026 | $2.00 | $10 | $0.20 | $4.00 | 2.7× the price |
| Claude Sonnet 5.5AnthropicCache reads were reduced from $0.20 to $0.10 per million tokens on 7 October 2026. | 28 Sept 2026 | $2.00 | $10 | $0.10 | $4.00 | 2.7× the price |
| Gemini 4 ArgonGoogle DeepMind · Limited accessAnnounced introductory rates for the upcoming broader launch. Regular rates will be $4 input and $20 output per million tokens; the introductory end date and batch discount are not published. | 30 Sept 2026 | $2.00 | $10 | $0.10 | $4.00 | 2.7× the price |
| GPT-6 SolOpenAIListed rates apply up to 272K input tokens. Above 272K, input costs $4, cached input $0.40 and output $15 per million tokens for the full request. | 22 Sept 2026 | $2.00 | $10 | $0.20 | $4.00 | 2.7× the price |
| GPT-6.1 SolOpenAIListed rates apply up to 272K input tokens. Above 272K, input costs $4, cached input $0.20 and output $15 per million tokens for the full request. | 29 Sept 2026 | $2.00 | $10 | $0.10 | $4.00 | 2.7× the price |
| Gemini 3.1 ProGoogle DeepMind · PreviewRates apply to prompts up to 200K tokens. Above 200K, input is $4, output $18 and cached input $0.40 per million tokens. Cache storage and grounding are billed separately. | 19 Feb 2026 | $2.00 | $12 | $0.20 | $4.50 | 3× the price |
| Claude Opus 5.5Anthropic | 22 Sept 2026 | $4.00 | $20 | $0.20 | $8.00 | 5.3× the price |
| GPT-5.6 SolOpenAIPrice before the GPT-6 launch, from OpenAI's announcement. OpenRouter now lists $2 / $10. | 9 Jul 2026 | $4.00 | $20 | – | $8.00 | 5.3× the price |
| Claude Opus 5Anthropic | 24 Jul 2026 | $5.00 | $25 | $0.50 | $10 | 6.7× the price |
| Claude Fable 5Anthropic | 9 Jun 2026 | $10 | $50 | $1.00 | $20 | 13× the price |
| Claude Fable 5.1Anthropic | 1 Sept 2026 | $10 | $50 | $0.25 | $20 | 13× the price |
| GPT-6 AstraOpenAI | 4 Sept 2026 | $10 | $50 | $1.00 | $20 | 13× the price |
GPT-6.1 Sol and GPT-6 Sol rates were checked on 8 October 2026 against OpenAI model documentation. USD per million tokens, checked 2 Oct 2026 via OpenRouter. Comparison uses the blended price (3 input tokens for every output token). Haiku 5.5 and Sonnet 5.5 prices were updated from Anthropic's 7 October 2026 announcement. Gemini and Grok prices were checked against official Google and xAI documentation on 8 October 2026.
Official benchmark results
Reported by Google DeepMind on 2 Sept 2026 · exact values
Exact published scores · Higher is better. Unreported results are omitted.
| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash | Claude Opus 5 | Claude Sonnet 5 | GPT-5.6 Sol | GPT-5.6 Terra |
|---|---|---|---|---|---|---|
| DeepSWE v1.1Long-horizon software engineering | 73.7% | 65.3% | 74% | 53.8% | 72.7% | 69.6% |
| GDPVal-AA v2Knowledge work · Elo | 1545 Elo | 1482 Elo | 1824 Elo | 1584 Elo | 1710 Elo | 1528 Elo |
| Vals Finance Agent v2Financial analyst tasks | 61.4% | 59% | 58.6% | 53.9% | 53.8% | 54.4% |
| Harvey's Legal Agent BenchmarkComplex legal workflows · All pass rate | 10% | 8.8% | 6.7% | 5% | 2.5% | 0.8% |
| Terminal-bench 2.1Agentic terminal coding | 89.4% | 85.8% | 89.1% | 80.4% | 88.8% | 87.4% |
| Terminal-bench 4.0General agent capabilities | 19.1% | 11.2% | 51.8% | 12.4% | 37.3% | 23.6% |
| GDP.PDFExpert PDF document comprehension · All pass rate | 35% | 34% | 37% | 28% | 40% | 29% |
| CharXiv ReasoningInformation synthesis from complex charts · No tools | 86.2% | 84.5% | 83.7% | 70.1% | 85.8% | 85.9% |
| LVBenchLong video understanding · (agentic) | 87.8% | 85.4% | 75.4% | 68.5% | 82.1% | 78.9% |
| LVBenchLong video understanding · (static) | 87.1% | – | – | – | – | – |
| HLE-VerifiedMultidisciplinary expert reasoning | 54.9% | 53.6% | 54.4% | 31% | 54.5% | 51.1% |
| OSWorld-2.0Agentic computer use · Partial score batch tool enabled | 59% | 50.6% | 75.4% | 42.6% | 62.6% | 50.2% |
| BioMysteryBenchBioinformatics research workflows · Human Solvable | 88.8% | 87.1% | 90.1% | 87.5% | 79.5% | 83.8% |
| BioMysteryBenchBioinformatics research workflows · Human Difficult | 56.5% | 43.5% | 49.4% | 34.1% | 44.7% | 49.4% |
| LABBench2Biology real-world research tasks | 86.2% | 82.1% | 84.2% | 80.1% | 82.1% | 81.2% |
View Google DeepMind's original charts



- Bold marks the best reported score in each row.
- Exact scores from the Google DeepMind model card. Gemini 3.7 Flash and GPT-5.6 Terra are benchmark comparisons only.
- LVBench reports separate agentic and static results for Gemini 3.8 Flash. Competitor cells span both rows in the source; they appear alongside the agentic result here, with no separate static values inferred. OSWorld-2.0 uses partial scoring with the batch tool enabled.
- These are Google-reported results; keep benchmark versions and evaluation harnesses in mind when comparing other launch tables.
Score against average cost per task · reported by Google DeepMind on 2 Sept 2026
Google's original comparison of long-horizon software engineering scores and task costs.
Google DeepMind has not published the numbers behind this chart, so the original image is shown instead. Open the announcement for the full context.

- The image reverses the cost axis: lower cost is to the right. Exact numeric task costs and per-effort point values were not published with this chart, so no reconstructed points are shown.
- The exact launch scores are available in the Gemini 3.8 Flash result table.