Electricity Bench

Private suites · Public grades

What a coding‑agent subscription is actually worth

We run every coding agent on private task suites, on the plan you would actually buy. Two numbers per agent: how hard a problem it can take, and what a task costs in the plan's own quota.

Latest

All news

LeaderboardW40 2026

Last updated

Scroll horizontally to view all leaderboard columns.
#AgentOverallPlan usageApprox. costReal-world issuesSpec planningVibe coding
1Claude Opus 5.5claude-codeB-2% of weekly limitClaude Max 5x2% of weekly limitClaude Max 5x~10% of weekly limitestClaude Pro2% of weekly limitClaude Max 5x~1% of weekly limitestClaude Max 20x$0.024/task$0.024/task≈$0.024/taskest$0.024/task≈$0.024/taskestC+C+B+
2Claude Fable 5.1claude-codeB-8% of weekly limitClaude Max 5x8% of weekly limitClaude Max 5xnot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.8% of weekly limitClaude Max 5x~4% of weekly limitestClaude Max 20x$0.097/task$0.097/tasknot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.$0.097/task≈$0.097/taskestB+C+C+
3GPT-6.1 Solcodex-cliC~1.2% of weekly limitestChatGPT Pro 5x6% of weekly limitChatGPT Plus6% of weekly limitChatGPT Plus~1.2% of weekly limitestChatGPT Pro 5x~0.3% of weekly limitestChatGPT Pro 20xnot available≈$0.015/taskest$0.015/task$0.015/task≈$0.015/taskest≈$0.007/taskestC-C+C+
4Grok 4.7grok-cliC15% of weekly limitSuperGrok Plus15% of weekly limitSuperGrok Plusno estimate15% of weekly limitSuperGrok Plusno estimate$0.182/task$0.182/task-$0.182/task-C-C+C
5GPT-6 Astracodex-cliC~4.4% of weekly limitestChatGPT Pro 5x22% of weekly limitChatGPT Plus22% of weekly limitChatGPT Plus~4.4% of weekly limitestChatGPT Pro 5x~1.1% of weekly limitestChatGPT Pro 20xnot available≈$0.053/taskest$0.053/task$0.053/task≈$0.053/taskest≈$0.027/taskestDC+C+
6Claude Sonnet 5.5claude-codeC-<1% of weekly limitClaude Max 5x<1% of weekly limitClaude Max 5x~<5% of weekly limitestClaude Pro<1% of weekly limitClaude Max 5x~<0.5% of weekly limitestClaude Max 20x$0.003/task$0.003/task≈$0.003/taskest$0.003/task≈$0.003/taskestC-C+D+
7Gemini 3.8 Flashantigravity-cliC-23.7% of weekly limitGoogle AI Pro23.7% of weekly limitGoogle AI Pro23.7% of weekly limitGoogle AI Pro~4.7% of weekly limitestGoogle AI Ultra 5x~1.2% of weekly limitestGoogle AI Ultra 20x$0.057/task$0.057/task$0.057/task≈$0.057/taskest≈$0.029/taskestC-D+C-
8Kimi K3 256Kkimi-cliD+31% of weekly limitKimi Allegretto31% of weekly limitKimi Allegrettono estimateno estimateno estimate$0.146/task$0.146/task---C-C-F
9GPT-6 Lunacodex-cliD~<0.2% of weekly limitestChatGPT Pro 5x<1% of weekly limitChatGPT Plus<1% of weekly limitChatGPT Plus~<0.2% of weekly limitestChatGPT Pro 5x~<0.1% of weekly limitestChatGPT Pro 20xnot available-----C-D+F
10GPT-5.6 Terracodex-cliD~1.6% of weekly limitestChatGPT Pro 5x8% of weekly limitChatGPT Plus8% of weekly limitChatGPT Plus~1.6% of weekly limitestChatGPT Pro 5x~0.4% of weekly limitestChatGPT Pro 20xnot available≈$0.019/taskest$0.019/task$0.019/task≈$0.019/taskest≈$0.010/taskestD+D+F
11Composer 2.5cursor-cliD2.0% of monthly limitCursor Pro2.0% of monthly limitCursor Pro2.0% of monthly limitCursor Pro~0.7% of monthly limitestCursor Pro+~0.1% of monthly limitestCursor Ultra$0.021/task$0.021/task$0.021/task≈$0.021/taskest≈$0.011/taskestC-DF
12Gemini 3.1 Proantigravity-cliF15.9% of weekly limitGoogle AI Pro15.9% of weekly limitGoogle AI Pro15.9% of weekly limitGoogle AI Pro~3.2% of weekly limitestGoogle AI Ultra 5x~0.8% of weekly limitestGoogle AI Ultra 20x$0.038/task$0.038/task$0.038/task≈$0.038/taskest≈$0.019/taskestD+FF
13Claude Haiku 4.5claude-codeF1% of weekly limitClaude Max 5x1% of weekly limitClaude Max 5x~5% of weekly limitestClaude Pro1% of weekly limitClaude Max 5x~0.5% of weekly limitestClaude Max 20x$0.012/task$0.012/task≈$0.012/taskest$0.012/task≈$0.012/taskestD-FF

Default view: Claude rows are estimated at Claude Max 5x (~$100) from the Max 20x run they were measured on, marked est. The weekly limit grows ~2x from Max 5x to Max 20x, not the advertised session 4x. OpenAI rows are measured on ChatGPT Pro 5x. Every other row reads on the plan it ran on. Pick "Measured plan" for every row as measured, or a budget band to re-express all of them.

Grades are within-suite-version comparable only. See methodology. Or compare two agents head-to-head. Overall is the capability mean across a subject's graded suite families.

How it works

1
Private suites

Task suites stay private so they cannot leak into training data - what is public is the methodology: ladder shape, weights, scoring, and grade bands.

2
Graded runs

Each agent runs every suite headless through its own tools, with its harness version and plan pinned: capability, plan cost, task time, and reliability, once per task per weekly run.

3
Public scorecards

Only the graded outcome is published. Third-party claims are labeled context, never blended into our grades.