Electricity Bench

Private suites · Public grades

What a coding-agent subscription is actually worth

We benchmark coding agents on the plans you would buy - Claude Code, Codex CLI, grok-cli and the rest, running real private task suites headless with their own tools - and answer the two questions that matter: how hard a problem can you hand this agent, and what does real work cost in the plan's own quota. The tasks stay private so nobody can train against them; the numbers come out in the open.

12agent scorecards
3suite families live

Leaderboard

Last updated

Grades look compressed on purpose: the suites are cut so the top rungs stay unsolved, and a Chere is a competitive agent rather than a poor one. The grade bands are published per suite.

Scroll horizontally to view all leaderboard columns.
#AgentOverallPlan usageApprox. costReal-world issuesSpec planningVibe codingLast run
1claude-fable-5claude-codeC12%/5hClaude Max 20xnot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.~48%/5hestClaude Max 5x12%/5hClaude Max 20x$0.041/tasknot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.≈$0.082/taskest$0.041/taskCCC
2claude-opus-5claude-codeC4%/5hClaude Max 20x~80%/5hestClaude Pro~16%/5hestClaude Max 5x4%/5hClaude Max 20x$0.014/task≈$0.027/taskest≈$0.027/taskest$0.014/taskCCC
3gpt-5.6-solcodex-cliC0.9%/1wChatGPT Plus0.9%/1wChatGPT Plus~0.2%/1westChatGPT Pro 5x~0.0%/1westChatGPT Pro 20x$0.010/task$0.010/task≈$0.010/taskest≈$0.005/taskestCCC
4grok-4.6grok-cliC5.4%/1wSuperGrok5.4%/1wSuperGrokno estimateno estimate$0.093/task$0.093/task--CCD
5gemini-3.7-flashantigravity-cliC7.3%/5hGoogle AI Pro7.3%/5hGoogle AI Pro~1.5%/5hestGoogle AI Ultra 5x~0.4%/5hestGoogle AI Ultra 20x$0.014/task$0.014/task≈$0.011/taskest≈$0.006/taskestDDB
6gpt-5.6-terracodex-cliD0.4%/1wChatGPT Plus0.4%/1wChatGPT Plus~0.1%/1westChatGPT Pro 5x~0.0%/1westChatGPT Pro 20x$0.005/task$0.005/task≈$0.005/taskest≈$0.002/taskestDCD
7kimi-k3-256kkimi-cliD51%/5hKimi Allegrettono estimateno estimateno estimate$0.034/task---CDD
8composer-2.5cursor-cliD0.6%/1mCursor Pro0.6%/1mCursor Pro~0.2%/1mestCursor Pro+~0.0%/1mestCursor Ultra$0.029/task$0.029/task≈$0.029/taskest≈$0.014/taskestCDD
9gpt-5.6-lunacodex-cliD0.3%/1wChatGPT Plus0.3%/1wChatGPT Plus~0.1%/1westChatGPT Pro 5x~0.0%/1westChatGPT Pro 20x$0.004/task$0.004/task≈$0.004/taskest≈$0.002/taskestDCF
10claude-sonnet-5claude-codeD3%/5hClaude Max 20x~60%/5hestClaude Pro~12%/5hestClaude Max 5x3%/5hClaude Max 20x$0.010/task≈$0.021/taskest≈$0.021/taskest$0.010/taskDDF
11claude-haiku-4-5claude-codeD1%/5hClaude Max 20x~20%/5hestClaude Pro~4%/5hestClaude Max 5x1%/5hClaude Max 20x$0.003/task≈$0.007/taskest≈$0.007/taskest$0.003/taskCFF
12gemini-3.1-proantigravity-cliF4.6%/5hGoogle AI Pro4.6%/5hGoogle AI Pro~0.9%/5hestGoogle AI Ultra 5x~0.2%/5hestGoogle AI Ultra 20x$0.009/task$0.009/task≈$0.007/taskest≈$0.004/taskestDFF

Grades are within-suite-version comparable only. See methodology. Or compare two agents head-to-head. Overall is the capability mean across a subject's graded suite families.

How it works

1
Private suites

Task suites stay private so they cannot leak into training data - what is public is the methodology: ladder shape, weights, scoring, and grade bands.

2
Graded runs

Each agent runs every suite headless through its own tools, with its harness version and plan pinned: capability, plan cost, task time, and reliability, with repeats for stability.

3
Public scorecards

Only the graded outcome is published. Third-party claims are labeled context, never blended into our grades.

Frequently asked questions

What is Electricity Bench?

An independent benchmark of AI coding subscriptions. We run 12 coding agents headless through their own tools on private task suites, grade the results under a public methodology, and publish every scorecard - including how much of each subscription plan a run actually consumes.

Which AI coding subscription is best?

It depends on what you value: peak capability, work per dollar, or ecosystem fit. The leaderboard ranks every graded agent by the tier of problem it can solve and shows what one run costs on its own plan; the two-minute which-subscription quiz turns your answers into a recommendation backed by these scorecards.

Why are the benchmark tasks private?

Public benchmark tasks leak into training data and get optimized against, which quietly inflates scores. Our suites stay private so every agent faces genuinely unseen problems; the methodology, grading bands, and full results are public instead.

How is plan usage measured?

Per suite run, from the provider's own usage meter where one exists (otherwise reconstructed from harness telemetry). Each figure is an absolute share of that agent's named plan - we never subtract one plan's share from another's, because different subscriptions are not the same unit. Cross-agent value comparisons use work per window and runs per plan-dollar instead.

Latest news

All news
weeklyClaude Fable 5 leads Week 34 as spec-planning debuts

A new suite tests whether agents can plan before they code, Kimi K3 joins the board, and Claude Fable 5 leads by having no weak spot.

releaseGemini 3.7 Flash: more capability for less quota

Gemini 3.7 Flash solves more real-world tasks than 3.6 Flash and uses less Google AI Pro quota - but it takes longer.

announcementNew suite: spec-planning - did the spec see the bug coming?

There is a new suite on the bench: spec-planning. It measures something none of our other suites do: can a coding agent look at a real codebase, read a big...