Electricity Bench

Private suites · Public grades

What a coding-agent subscription is actually worth

We benchmark coding agents on the plans you would buy - Claude Code, Codex CLI, grok-cli and the rest, running real private task suites headless with their own tools - and answer the two questions that matter: how hard a problem can you hand this agent, and what does real work cost in the plan's own quota. The tasks stay private so nobody can train against them; the numbers come out in the open.

10agents graded
8suite families live
0vendor claims counted
Every coding agent’s scorecard, the day it lands.

Leaderboard

Last updated

Scroll horizontally to view all leaderboard columns.
#AgentPlan usage / runOverallLast run
1claude-code + claude-fable-5no usage data-pending
2claude-code + claude-haiku-4-52.0% of plan · 50.0 runs/5hClaude Max 5x · provider meter-pending
3claude-code + claude-opus-4-82.0% of plan · 50.0 runs/5hClaude Max 5x · provider meter-pending
4claude-code + claude-sonnet-51.0% of plan · 100.0 runs/5hClaude Max 5x · provider meter-pending
5codex-cli + gpt-5.6-luna0.1% of plan · 1685.4 runs/weekChatGPT Plus-pending
6codex-cli + gpt-5.6-sol0.1% of plan · 1526.5 runs/weekChatGPT Plus-pending
7codex-cli + gpt-5.6-terra0.1% of plan · 1527.2 runs/weekChatGPT Plus-pending
8cursor-cli + composer-2.56.4% of plan · 15.6 runs/monthCursor Free-pending
9grok-cli + grok-4.50.3% of plan · 317.4 runs/weekSuperGrok-pending
10kimi-cli + kimi-k33.0% of plan · 33.3 runs/5hKimi Allegretto · provider meter-pending

Grades are within-suite-version comparable only. See methodology. Or compare two agents head-to-head. Overall grades appear once a subject is graded on two distinct suite families.

Table stakes

Saturated floor checks: every capable harness passes these, so they no longer carry the headline. Still graded - a subject failing one is news.

AgentStructured extraction
claude-code + claude-fable-5Av1
claude-code + claude-haiku-4-5Av1
claude-code + claude-opus-4-8Av1
claude-code + claude-sonnet-5Av1
codex-cli + gpt-5.6-lunaAv1
codex-cli + gpt-5.6-solAv1
codex-cli + gpt-5.6-terraAv1
cursor-cli + composer-2.5Av1
grok-cli + grok-4.5Av1
kimi-cli + kimi-k3Av1

How it works

1
Private suites

Niche task suites - agentic coding, extraction, automation - stay private so they cannot leak into training data. See a public sample task.

2
Graded runs

Each agent runs every suite headless through its own tools, with its harness version and plan pinned: capability, plan cost, latency, and reliability, with repeats for stability.

3
Public scorecards

Only the graded outcome is published. Third-party claims are labeled context, never blended into our grades.

Latest news

All news
releaseIntroducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Google is introducing new Gemini models, including Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber. These are newly shipped…

announcementIntroducing Electricity Bench

Electricity Bench grades new AI models on private, niche task suites and publishes the results here: scorecards, an AI-tools directory, and news.

postWhy honest model numbers are hard

Publishing a fair number for a model is harder than it looks, and most of the difficulty is invisible in the final figure.