Electricity Bench

2026-W32: Real-world issues

Week 32 of 2026 graded 11 coding agents on real-world issues. Gemini 3.8 Flash leads with C-. Grades compare only within the same suite version, and plan usage is each agent's share of its own subscription. Every agent ran the same suite cut headless through its own tools.Measured

real-world@v3 - the suite key these grades compare within

Published suite snapshot for 11 graded agents, current through Aug 5, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Real-world issues methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
Gemini 3.8 Flashantigravity-cliC-60.0%tier 4 · partial tier 52m 32s100%19.6%/1wGoogle AI Pro
Claude Fable 5.1claude-codeC-60.0%tier 4 · partial tier 52m 49s100%122%/5hClaude Max 5x
Claude Opus 5.5claude-codeC-60.0%tier 4 · partial tier 53m 43s100%64%/5hClaude Max 5x
Grok 4.6grok-cliC-60.0%tier 4 · partial tier 54m 7s100%2.4%/1wSuperGrok Plus
Composer 2.5cursor-cliC-56.7%tier 2 · partial tier 51m 57s100%1.7%/1mCursor Pro
Gemini 3.1 Proantigravity-cliD46.7%tier 4 · partial tier 52m 31s100%16.8%/1wGoogle AI Pro
GPT-6 Lunacodex-cliD46.7%tier 4 · partial tier 51m 35s100%0.7%/1wChatGPT Plus
GPT-6 Solcodex-cliD46.7%tier 4 · partial tier 53m 31s100%11.4%/1wChatGPT Plus
Claude Sonnet 5claude-codeF33.3%tier 3 · partial tier 53m 48s100%35%/5hClaude Max 5x
Claude Haiku 4.5claude-codeF30.0%tier 2 · partial tier 53m 9s100%17%/5hClaude Max 5x
GPT-5.6 Terracodex-cliF27.5%none1m 26s100%0.9%/1wChatGPT Pro 5x

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.