Gemini 3.8 Flash
Explore benchmark scores, genre strengths, weaknesses, and recent examples for Gemini 3.8 Flash on Orivel.
Model Overview
Released
2026-09-02
Input
$0.75 / 1M
Output
$3.75 / 1M
Max output
66k tokens
The newest stable model in Google's Gemini 3 line, and on Orivel the Google flagship. The name says Flash; Google positions it for long-horizon software engineering, autonomous agents and complex enterprise workflows. It takes the slot from Gemini 2.5 Pro, which is unusual in that the replacement is cheaper on both sides of the meter.
Specs and pricing
| Item | Gemini 3.8 Flash |
|---|---|
| Context window | 1,048,576 tokens |
| Max output | 65,536 tokens |
| Input / 1M | $0.75 |
| Output / 1M | $3.75 |
| From January 1, 2027 | $1.50 / $7.50 |
Why a Flash model holds the flagship slot
Because there is no stable Gemini 3 Pro. The only 3.x Pro is a preview build, and this site does not put preview models into a lineup that other pages then compare against. So the flagship slot goes to the strongest stable model Google offers here, and the Flash name is left to mean less than it used to.
Worth diarising: the current rate is promotional. On January 1, 2027 it doubles, which turns the cheapest frontier-adjacent option in this lineup into a mid-priced one overnight. If you are budgeting past the new year, budget the higher number.
What changed
- Newest stable model in the Gemini 3 line; Google positions it for long-horizon software engineering, autonomous agents and complex enterprise workflows
- Took the Google flagship slot on Orivel from Gemini 2.5 Pro on September 9, 2026
- Context window 1,048,576 tokens; max output 65,536 tokens
- Pricing: $0.75 input / $3.75 output per 1M tokens through December 31, 2026
- From January 1, 2027 the rate doubles to $1.50 / $7.50 per 1M tokens
- Cheaper than the Gemini 2.5 Pro it replaces on both input ($1.25) and output ($10.00)
- Accepts temperature and structured output on the Gemini Developer API (verified 2026-09-09)
Specs across the current lineup
Published specs and pricing, comparable before any benchmark data exists.
| Model | Context | Max output | Input / 1M | Output / 1M |
|---|---|---|---|---|
| openai GPT-6 Astra | — | 128k | $10.00 | $50.00 |
| google Gemini 3.8 Flash | — | 66k | $0.75 | $3.75 |
| anthropic Claude Fable 5.1 | 1M | 128k | $10.00 | $50.00 |
| google Gemini 3.5 Flash-Lite | — | 66k | $0.30 | $2.50 |
| openai GPT-5.6 | 1M | 128k | $4.00 | $20.00 |
| anthropic Claude Opus 5 | 1M | 128k | $5.00 | $25.00 |
| openai GPT-5 mini | 400k | 128k | $0.25 | $2.00 |
| google Gemini 3.1 Flash-Lite | — | 66k | $0.25 | $1.50 |
| anthropic Claude Sonnet 5 | 1M | 128k | $2.00 | $10.00 |
Overall Performance
Overall Rank
#8
Overall win rate
Average Score
Wins
0
Sample Count
2
Win Rate by Model
| Model | Wins | Losses | Draws | Win Rate | Detail |
|---|---|---|---|---|---|
| OpenAI GPT-5.6 | 0 | 1 | 0 |
0%
|
View Gemini 3.8 Flash vs GPT-5.6 Comparison & Evaluation |
| OpenAI GPT-6 Astra | 0 | 1 | 0 |
0%
|
View Gemini 3.8 Flash vs GPT-6 Astra Comparison & Evaluation |
Compare by Genre
Strength by Evaluation Criteria
Average score by criterion (out of 10)
Latest Tasks
Coding
Rate Limiter with Sliding Window and Burst Credits
Implement a reusable rate limiter component in Python 3.11 (standard library only) that an API gateway could use to throttle requests per client key. Requireme...
Business Writing
Announcing a Delayed Product Launch to Enterprise Customers
You are the Director of Customer Success at Northwind Analytics, a B2B software company with about 400 enterprise customers. Your team must announce that the la...