What a crazy week behind us. We started off with the release of Claude Fable 5.1, then came Gemini 3.8 and finally GPT-6 Astra. The week 2026-W37 grading run is the first weekly run to grade them side by side, across real-world@v3, spec-planning@v1 and vibe-coding@v1.
What the results show
- Claude Fable 5.1 wins the week over GPT-6 Astra. Fable is the only agent that stayed solid across all three suites and it holds the top spot on the week's hardest tasks. The two run at about the same speed, a few minutes per task at the median. Usage is where they part: at the same monthly price Astra gets through a grading pass on about half the quota Fable needs.
- Gemini 3.8 Flash burned heavy again. For the second run in a row one pass of the real-world issues takes about a fifth of its Google AI Pro week, roughly double what 3.7 Flash used to spend. That has us rethinking what to reach for as the cheap driver.
- Kimi K3's usage is insane. One full grading pass costs it 180% of a five-hour window on the $39 Kimi Allegretto plan, nearly two windows for what the bench asks of every agent, which comes out at 38% of its week. If the Claude and Codex plans cost the same $39, the same pass would take Claude Fable 5.1 about a third of its week and GPT-6 Astra about an eighth.
The week goes to Claude Fable 5.1, with GPT-6 Astra right behind. The part we keep thinking about is the cheap model. From experience, several models got worse after the release of a newer version. Our suspicion is that providers trim a model to save money once its successor ships, through heavier quantization or a lower reasoning effort. We cannot prove this, so read it as a guess. With that in mind we have to find a successor for Gemini 3.7 as our driver for smaller tasks (our take).
Grades are comparable within the same suite version. Explore the complete Week 37 results or follow each agent over time on Trends.