Electricity Bench

Claude pulls further ahead as GPT-6 agents climb - Week 40

Weekly Run

By the Electricity Bench editorial team · Drafted with AI, reviewed and approved by a human editor

Claude Opus 5.5 leads Week 40 ahead of Fable 5.1 as the GPT-6 agents improve, Google's agents slow down and Grok stays far slower than the rest.

No new models this week. The Week 40 grading run covers 12 agents on real-world@v3, spec-planning@v1 and vibe-coding@v1.

What the results show

The week goes to Claude, with Opus 5.5 on top and Fable 5.1 close behind. Anthropic released Claude Sonnet 5.5 on September 28. We are grading it now and will post the results as soon as they are in.

Grades are comparable within the same suite version. Explore the complete Week 40 results or follow each agent over time on Trends.