Private suites · Public grades
What a coding-agent subscription is actually worth
We benchmark coding agents on the plans you would buy - Claude Code, Codex CLI, grok-cli and the rest, running real private task suites headless with their own tools - and answer the two questions that matter: how hard a problem can you hand this agent, and what does real work cost in the plan's own quota. The tasks stay private so nobody can train against them; the numbers come out in the open.
Not sure which plan fits you? Take the five-question quiz - it recommends a plan and links the scorecards behind it.
Leaderboard
Last updated
| # | Agent | Plan usage / run | Overall | Last run |
|---|---|---|---|---|
| 1 | claude-code + claude-fable-5 | no usage data | -pending | |
| 2 | claude-code + claude-haiku-4-5 | 2.0% of plan · 50.0 runs/5hClaude Max 5x · provider meter | -pending | |
| 3 | claude-code + claude-opus-4-8 | 2.0% of plan · 50.0 runs/5hClaude Max 5x · provider meter | -pending | |
| 4 | claude-code + claude-sonnet-5 | 1.0% of plan · 100.0 runs/5hClaude Max 5x · provider meter | -pending | |
| 5 | codex-cli + gpt-5.6-luna | 0.1% of plan · 1685.4 runs/weekChatGPT Plus | -pending | |
| 6 | codex-cli + gpt-5.6-sol | 0.1% of plan · 1526.5 runs/weekChatGPT Plus | -pending | |
| 7 | codex-cli + gpt-5.6-terra | 0.1% of plan · 1527.2 runs/weekChatGPT Plus | -pending | |
| 8 | cursor-cli + composer-2.5 | 6.4% of plan · 15.6 runs/monthCursor Free | -pending | |
| 9 | grok-cli + grok-4.5 | 0.3% of plan · 317.4 runs/weekSuperGrok | -pending | |
| 10 | kimi-cli + kimi-k3 | 3.0% of plan · 33.3 runs/5hKimi Allegretto · provider meter | -pending |
Grades are within-suite-version comparable only. See methodology. Or compare two agents head-to-head. Overall grades appear once a subject is graded on two distinct suite families.
Table stakes
Saturated floor checks: every capable harness passes these, so they no longer carry the headline. Still graded - a subject failing one is news.
How it works
Niche task suites - agentic coding, extraction, automation - stay private so they cannot leak into training data. See a public sample task.
Each agent runs every suite headless through its own tools, with its harness version and plan pinned: capability, plan cost, latency, and reliability, with repeats for stability.
Only the graded outcome is published. Third-party claims are labeled context, never blended into our grades.
Latest news
All newsGoogle is introducing new Gemini models, including Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. These are newly shipped…
Introducing Electricity BenchElectricity Bench grades new AI models on private, niche task suites and publishes the results here: scorecards, an AI-tools directory, and news.
Why honest model numbers are hardPublishing a fair number for a model is harder than it looks, and most of the difficulty is invisible in the final figure.