Private suites · Public grades
What a coding-agent subscription is actually worth
We benchmark coding agents on the plans you would buy - Claude Code, Codex CLI, grok-cli and the rest, running real private task suites headless with their own tools - and answer the two questions that matter: how hard a problem can you hand this agent, and what does real work cost in the plan's own quota. The tasks stay private so nobody can train against them; the numbers come out in the open.
Not sure which plan fits you? Take the five-question quiz - it recommends a plan and links the scorecards behind it.
Leaderboard
Last updated
Grades look compressed on purpose: the suites are cut so the top rungs stay unsolved, and a Chere is a competitive agent rather than a poor one. The grade bands are published per suite.
| # | Agent | Overall | Plan usage | Approx. cost | Real-world issues | Spec planning | Vibe coding | Last run |
|---|---|---|---|---|---|---|---|---|
| 1 | claude-fable-5claude-code | C | 12%/5hClaude Max 20xnot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.~48%/5hestClaude Max 5x12%/5hClaude Max 20x | $0.041/tasknot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.≈$0.082/taskest$0.041/task | C | C | C | |
| 2 | claude-opus-5claude-code | C | 4%/5hClaude Max 20x~80%/5hestClaude Pro~16%/5hestClaude Max 5x4%/5hClaude Max 20x | $0.014/task≈$0.027/taskest≈$0.027/taskest$0.014/task | C | C | C | |
| 3 | gpt-5.6-solcodex-cli | C | 0.9%/1wChatGPT Plus0.9%/1wChatGPT Plus~0.2%/1westChatGPT Pro 5x~0.0%/1westChatGPT Pro 20x | $0.010/task$0.010/task≈$0.010/taskest≈$0.005/taskest | C | C | C | |
| 4 | grok-4.6grok-cli | C | 5.4%/1wSuperGrok5.4%/1wSuperGrokno estimateno estimate | $0.093/task$0.093/task-- | C | C | D | |
| 5 | gemini-3.7-flashantigravity-cli | C | 7.3%/5hGoogle AI Pro7.3%/5hGoogle AI Pro~1.5%/5hestGoogle AI Ultra 5x~0.4%/5hestGoogle AI Ultra 20x | $0.014/task$0.014/task≈$0.011/taskest≈$0.006/taskest | D | D | B | |
| 6 | gpt-5.6-terracodex-cli | D | 0.4%/1wChatGPT Plus0.4%/1wChatGPT Plus~0.1%/1westChatGPT Pro 5x~0.0%/1westChatGPT Pro 20x | $0.005/task$0.005/task≈$0.005/taskest≈$0.002/taskest | D | C | D | |
| 7 | kimi-k3-256kkimi-cli | D | 51%/5hKimi Allegrettono estimateno estimateno estimate | $0.034/task--- | C | D | D | |
| 8 | composer-2.5cursor-cli | D | 0.6%/1mCursor Pro0.6%/1mCursor Pro~0.2%/1mestCursor Pro+~0.0%/1mestCursor Ultra | $0.029/task$0.029/task≈$0.029/taskest≈$0.014/taskest | C | D | D | |
| 9 | gpt-5.6-lunacodex-cli | D | 0.3%/1wChatGPT Plus0.3%/1wChatGPT Plus~0.1%/1westChatGPT Pro 5x~0.0%/1westChatGPT Pro 20x | $0.004/task$0.004/task≈$0.004/taskest≈$0.002/taskest | D | C | F | |
| 10 | claude-sonnet-5claude-code | D | 3%/5hClaude Max 20x~60%/5hestClaude Pro~12%/5hestClaude Max 5x3%/5hClaude Max 20x | $0.010/task≈$0.021/taskest≈$0.021/taskest$0.010/task | D | D | F | |
| 11 | claude-haiku-4-5claude-code | D | 1%/5hClaude Max 20x~20%/5hestClaude Pro~4%/5hestClaude Max 5x1%/5hClaude Max 20x | $0.003/task≈$0.007/taskest≈$0.007/taskest$0.003/task | C | F | F | |
| 12 | gemini-3.1-proantigravity-cli | F | 4.6%/5hGoogle AI Pro4.6%/5hGoogle AI Pro~0.9%/5hestGoogle AI Ultra 5x~0.2%/5hestGoogle AI Ultra 20x | $0.009/task$0.009/task≈$0.007/taskest≈$0.004/taskest | D | F | F |
Budget view: rows whose measured plan sits in the picked band keep their measured figure; the rest are scaled from the measured plan by the provider's own advertised tier multiples (marked est) - approximate, never measured. Providers that publish no multiple get no estimate. Our paired Claude Max 5x/20x runs measured ~5.7x where 4x is advertised, so treat estimates as rough. A model the tier does not sell at all - Fable on Claude Pro - reads not available rather than an estimate.
Grades are within-suite-version comparable only. See methodology. Or compare two agents head-to-head. Overall is the capability mean across a subject's graded suite families.
How it works
Task suites stay private so they cannot leak into training data - what is public is the methodology: ladder shape, weights, scoring, and grade bands.
Each agent runs every suite headless through its own tools, with its harness version and plan pinned: capability, plan cost, task time, and reliability, with repeats for stability.
Only the graded outcome is published. Third-party claims are labeled context, never blended into our grades.
Frequently asked questions
What is Electricity Bench?
An independent benchmark of AI coding subscriptions. We run 12 coding agents headless through their own tools on private task suites, grade the results under a public methodology, and publish every scorecard - including how much of each subscription plan a run actually consumes.
Which AI coding subscription is best?
It depends on what you value: peak capability, work per dollar, or ecosystem fit. The leaderboard ranks every graded agent by the tier of problem it can solve and shows what one run costs on its own plan; the two-minute which-subscription quiz turns your answers into a recommendation backed by these scorecards.
Why are the benchmark tasks private?
Public benchmark tasks leak into training data and get optimized against, which quietly inflates scores. Our suites stay private so every agent faces genuinely unseen problems; the methodology, grading bands, and full results are public instead.
How is plan usage measured?
Per suite run, from the provider's own usage meter where one exists (otherwise reconstructed from harness telemetry). Each figure is an absolute share of that agent's named plan - we never subtract one plan's share from another's, because different subscriptions are not the same unit. Cross-agent value comparisons use work per window and runs per plan-dollar instead.
Latest news
All newsA new suite tests whether agents can plan before they code, Kimi K3 joins the board, and Claude Fable 5 leads by having no weak spot.
Gemini 3.7 Flash: more capability for less quotaGemini 3.7 Flash solves more real-world tasks than 3.6 Flash and uses less Google AI Pro quota - but it takes longer.
New suite: spec-planning - did the spec see the bug coming?There is a new suite on the bench: spec-planning. It measures something none of our other suites do: can a coding agent look at a real codebase, read a big...