Electricity Bench

Agent scorecards

One scorecard per graded agent - a harness plus the model it drives, run on the plan you would buy. Ordered by the leaderboard capability ranking.

Cclaude-fable-5

#1 on the leaderboard · claude-code · Claude Max 20x

Cgemini-3.6-flash

#2 on the leaderboard · antigravity-cli · Google AI Pro

Cgrok-4.5

#3 on the leaderboard · grok-cli · SuperGrok

Cgpt-5.6-sol

#4 on the leaderboard · codex-cli · ChatGPT Plus

Ccomposer-2.5

#5 on the leaderboard · cursor-cli · Cursor Pro

Cclaude-opus-5

#6 on the leaderboard · claude-code · Claude Max 20x

Dgpt-5.6-luna

#7 on the leaderboard · codex-cli · ChatGPT Plus

Dgpt-5.6-terra

#8 on the leaderboard · codex-cli · ChatGPT Plus

Dclaude-sonnet-5

#9 on the leaderboard · claude-code · Claude Max 20x

Dgemini-3.1-pro

#10 on the leaderboard · antigravity-cli · Google AI Pro

Fclaude-haiku-4-5

#11 on the leaderboard · claude-code · Claude Max 20x

Grades are within-suite-version comparable only - see the methodology, or compare two agents head-to-head.