Electricity Bench

2026-W40: Real-world issues

Week 40 of 2026 graded 12 coding agents on real-world issues. Claude Fable 5.1 leads with B+. The biggest move since W39 2026 is Claude Fable 5.1, D+ to B+. Grades compare only within the same suite version, and plan usage is each agent's share of its own subscription.Measured

real-world@v3 - the suite key these grades compare within

Published suite snapshot for 12 graded agents, current through Sep 28, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Real-world issues methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
Claude Fable 5.1claude-codeB+86.7%tier 4 · partial tier 53m 38s100%64%/5hClaude Max 5x
Claude Opus 5.5claude-codeC+73.3%tier 4 · partial tier 51m 17s100%11%/5hClaude Max 5x
Gemini 3.8 Flashantigravity-cliC-60.0%tier 4 · partial tier 55m 21s100%21.9%/1wGoogle AI Pro
Kimi K3 256Kkimi-cliC-60.0%tier 4 · partial tier 53m 31s100%104%/5hKimi Allegretto
Grok 4.7grok-cliC-58.3%tier 1 · partial tier 58m 13s100%5%/1wSuperGrok Plus
GPT-6 Lunacodex-cliC-56.7%tier 2 · partial tier 51m 4s100%0%/1wChatGPT Plus
Composer 2.5cursor-cliC-56.7%tier 1 · partial tier 51m 34s100%1.6%/1mCursor Pro
GPT-6 Solcodex-cliC-55.0%tier 1 · partial tier 51m 41s100%5%/1wChatGPT Plus
Gemini 3.1 Proantigravity-cliD+53.3%tier 3 · partial tier 54m 53s100%14.8%/1wGoogle AI Pro
Claude Sonnet 5claude-codeD+53.3%tier 3 · partial tier 53m 35s100%12%/5hClaude Max 5x
GPT-6 Astracodex-cliD43.3%tier 2 · partial tier 52m 13s100%17%/1wChatGPT Plus
Claude Haiku 4.5claude-codeD-36.7%tier 2 · partial tier 53m 8s100%8%/5hClaude Max 5x

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.