Google released gemini-3.8-flash on 2026-09-02, three weeks after 3.7 Flash. We graded it the same day on Antigravity CLI, on the Google AI Pro plan at effort medium. It replaces Gemini 3.7 Flash on the bench. In short: it does not solve more, and it costs more to run.
What the results show
- No improvement. Gemini 3.8 Flash does not solve more issues on the easier tasks than 3.7 Flash did, and it is not an improvement on the planning tasks either (Week 36 results).
- It uses more quota. A run takes about 20% of the Google AI Pro weekly window, up from 8% for 3.7 Flash.
- It is slower. The same tasks take about twice as long (median 4m 14s, up from 2m 38s).
This feels like a downgrade. 3.7 Flash was crazy cheap for what it solved, which is what made it worth having on a Google AI Pro plan. 3.8 Flash keeps the grades and drops the cheap part. Grok 4.6 did the same to Grok 4.5: no gain in capability, and a move from cheap middle of the pack to normal price middle of the pack, where it gets beaten by the models it now costs as much as. Flash has made the same move, and in its new range it gets beaten by pretty much everything. We do not recommend it. Staying on 3.7 Flash makes the most sense for now. That may change as the harness improves, and we will keep grading it every week.
Grades are comparable within the same suite version. See the full Gemini 3.8 Flash scorecard or follow the family over time on Trends.