Electricity Bench

2026-W35: Real-world issues

real-world@v3 - the suite key these grades compare within

Published suite snapshot for 12 graded agents, current through Aug 24, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Real-world issues methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
claude-fable-5claude-codeC73.3%tier 4 · partial tier 53m 29s100%32%/5hClaude Max 20x
gemini-3.1-proantigravity-cliC66.7%tier 3 · partial tier 53m 34s100%9.5%/1wGoogle AI Pro
gemini-3.7-flashantigravity-cliC60.0%tier 4 · partial tier 53m 5s100%8.5%/1wGoogle AI Pro
claude-opus-5claude-codeC60.0%tier 4 · partial tier 53m 49s100%11%/5hClaude Max 20x
kimi-k3-256kkimi-cliD53.3%tier 3 · partial tier 53m 49s100%144%/5hKimi Allegretto
composer-2.5cursor-cliD50.0%tier 2 · partial tier 52m 10s100%1.8%/1mCursor Pro
claude-sonnet-5claude-codeD46.7%tier 4 · partial tier 52m 59s100%5%/5hClaude Max 20x
gpt-5.6-lunacodex-cliD46.7%tier 4 · partial tier 52m 5s100%3.4%/1wChatGPT Plus
claude-haiku-4-5claude-codeD43.3%tier 2 · partial tier 53m 33s100%4%/5hClaude Max 20x
gpt-5.6-solcodex-cliD43.3%tier 2 · partial tier 52m 53s100%3.8%/1wChatGPT Plus
gpt-5.6-terracodex-cliD43.3%tier 2 · partial tier 51m 40s100%2.5%/1wChatGPT Plus
grok-4.6grok-cliD40.0%tier 3 · partial tier 55m 34s100%3.3%/1wSuperGrok

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.