Electricity Bench

2026-W38: Real-world issues

Week 38 of 2026 graded 13 coding agents on real-world issues. Claude Opus 5 leads with C+. The biggest move since W37 2026 is Claude Opus 5, D+ to C+. Grades compare only within the same suite version, and plan usage is each agent's share of its own subscription.Measured

real-world@v3 - the suite key these grades compare within

Published suite snapshot for 13 graded agents, current through Sep 14, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Real-world issues methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
Claude Opus 5claude-codeC+73.3%tier 4 · partial tier 54m 3s100%9%/5hClaude Max 20x
Claude Fable 5.1claude-codeC-60.0%tier 4 · partial tier 51m 48s100%25%/5hClaude Max 20x
GPT-5.6 Solcodex-cliC-60.0%tier 4 · partial tier 53m 11s100%2%/1wChatGPT Pro 5x
Grok 4.6grok-cliC-60.0%tier 4 · partial tier 57m 19s100%5.9%/1wSuperGrok
Kimi K3 256Kkimi-cliC-60.0%tier 4 · partial tier 51m 27s100%114%/5hKimi Allegretto
Claude Sonnet 5claude-codeC-58.3%tier 1 · partial tier 54m 17s100%5%/5hClaude Max 20x
Gemini 3.8 Flashantigravity-cliD+53.3%tier 3 · partial tier 54m 51s100%18.4%/1wGoogle AI Pro
GPT-5.6 Terracodex-cliD+53.3%tier 3 · partial tier 52m 14s100%1%/1wChatGPT Pro 5x
Composer 2.5cursor-cliD+50.0%tier 2 · partial tier 51m 26s100%1.8%/1mCursor Pro
GPT-5.6 Lunacodex-cliD46.7%tier 4 · partial tier 51m 50s100%0%/1wChatGPT Pro 5x
Claude Haiku 4.5claude-codeD43.3%tier 2 · partial tier 53m 52s100%3%/5hClaude Max 20x
GPT-6 Astracodex-cliD43.3%tier 2 · partial tier 52m 16s100%5%/1wChatGPT Pro 5x
Gemini 3.1 Proantigravity-cliF33.3%tier 3 · partial tier 53m 30s100%11.2%/1wGoogle AI Pro

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.