No new models this week. The Week 40 grading run covers 12 agents on real-world@v3, spec-planning@v1 and vibe-coding@v1.
What the results show
- Claude Opus 5.5 stays the agent to beat. Claude got better again. Claude Opus 5.5 keeps the top of the board. Claude Fable 5.1 is right behind it after solving far more of the real issues than last week. Claude Code shipped several updates since then. Our guess is that part of the gain comes from the harness and not only from the models, but one week cannot tell the two apart.
- The ChatGPT agents are getting better. GPT-6 Astra and GPT-6 Sol both moved up a step on the same tasks as last week. They now sit fourth and fifth on the board. GPT-6 Luna held its place.
- Google's agents got slower. Gemini 3.1 Pro took about a third longer per task than last week. Gemini 3.8 Flash solved more of the real issues and moved up, but it also took longer. It is now the second slowest agent on the bench.
- Grok 4.7 is still three to four times slower than the rest. Grok 4.7 moved up to third on the board. But it needs about ten minutes per task, while Claude Opus 5.5 and the GPT-6 agents need under three.
The week goes to Claude, with Opus 5.5 on top and Fable 5.1 close behind. Anthropic released Claude Sonnet 5.5 on September 28. We are grading it now and will post the results as soon as they are in.
Grades are comparable within the same suite version. Explore the complete Week 40 results or follow each agent over time on Trends.