Electricity Bench

2026-W37: Real-world issues

real-world@v3 - the suite key these grades compare within

Published suite snapshot for 13 graded agents, current through Sep 7, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Real-world issues methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
Gemini 3.8 Flashantigravity-cliC73.3%tier 4 · partial tier 55m 42s100%21.6%/1wGoogle AI Pro
Claude Fable 5.1claude-codeC73.3%tier 4 · partial tier 52m 2s100%27%/5hClaude Max 20x
Composer 2.5cursor-cliC60.0%tier 4 · partial tier 51m 16s100%1.7%/1mCursor Pro
Grok 4.6grok-cliC60.0%tier 4 · partial tier 56m 14s100%3.0%/1wSuperGrok
Kimi K3 256Kkimi-cliC60.0%tier 4 · partial tier 53m 12s100%123%/5hKimi Allegretto
Claude Sonnet 5claude-codeC56.7%tier 2 · partial tier 53m 14s100%5%/5hClaude Max 20x
GPT-5.6 Solcodex-cliC56.7%tier 2 · partial tier 52m 27s100%2%/1wChatGPT Pro $100
GPT-5.6 Terracodex-cliD53.3%tier 3 · partial tier 52m 30s100%2%/1wChatGPT Pro $100
GPT-5.6 Lunacodex-cliD52.5%none2m 12s100%0%/1wChatGPT Pro $100
Claude Haiku 4.5claude-codeD51.7%tier 1 · partial tier 53m 35s100%4%/5hClaude Max 20x
Claude Opus 5claude-codeD50.0%tier 2 · partial tier 55m 8s100%10%/5hClaude Max 20x
Gemini 3.1 Proantigravity-cliD46.7%tier 3 · partial tier 53m 27s100%12.4%/1wGoogle AI Pro
GPT-6 Astracodex-cliD43.3%tier 2 · partial tier 52m 18s100%4%/1wChatGPT Pro $100

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.