No new models this week. The Week 39 grading run covers the same 13 agents on real-world@v3, spec-planning@v1 and vibe-coding@v1.
What the results show
- Claude Opus 5 is the agent to beat. Claude Opus 5 moves to the top of the board and answered about twice as fast as last week. That matches our own use. Opus has become the Claude model we reach for when coding, because Claude Fable 5.1 burns too much usage for everyday work. Fable still writes the best plans on the bench.
- Max 20x buys four times the session but only twice the week. On Claude Max 5x, the $100 plan, one pass over the real issues takes Fable's whole five-hour window and about a fifth of it for Opus. On Max 20x the same work took Fable a quarter of the window. Anthropic's 5x and 20x describe the five-hour window. The weekly limit only doubles between the two plans.
- Most of the field slipped a little. Seven of 13 agents dropped a step overall on the same tasks as last week, among them GPT-6 Astra, GPT-5.6 Sol, Grok 4.6 and Fable. None moved up. Our guess, and it is only a guess, is that providers make a model cheaper to serve once it has been out for a while, with lower effort or heavier quantization, and it costs a little quality.
- Kimi K3 is slow again. Last week's speed-up turned out to be a fluke. Kimi K3 is back to about four minutes per real issue, with the same heavy usage.
The week goes to Claude Opus 5, with Fable 5.1 still the one to pick for planning. There is a lot of talk about several providers releasing new models this week, so the board may change soon. We will grade each one as it lands.
Grades are comparable within the same suite version. Explore the complete Week 39 results or follow each agent over time on Trends.