Electricity Bench

Agent scorecards

One scorecard per graded agent - a harness plus the model it drives, run on the plan you would buy. Ordered by the leaderboard capability ranking.

C+Claude Opus 5.5

#1 on the leaderboard · claude-code · Claude Max 5x

C-Claude Fable 5.1

#2 on the leaderboard · claude-code · Claude Max 5x

C-Grok 4.7

#3 on the leaderboard · grok-cli · SuperGrok Plus

C-GPT-6 Astra

#4 on the leaderboard · codex-cli · ChatGPT Pro 5x

D+GPT-6 Sol

#5 on the leaderboard · codex-cli · ChatGPT Plus

D+GPT-5.6 Terra

#6 on the leaderboard · codex-cli · ChatGPT Pro 5x

DClaude Sonnet 5

#7 on the leaderboard · claude-code · Claude Max 5x

DGPT-6 Luna

#8 on the leaderboard · codex-cli · ChatGPT Plus

DKimi K3 256K

#9 on the leaderboard · kimi-cli · Kimi Allegretto

DGemini 3.8 Flash

#10 on the leaderboard · antigravity-cli · Google AI Pro

D-Composer 2.5

#11 on the leaderboard · cursor-cli · Cursor Pro

FGemini 3.1 Pro

#12 on the leaderboard · antigravity-cli · Google AI Pro

FClaude Haiku 4.5

#13 on the leaderboard · claude-code · Claude Max 5x

Grades are within-suite-version comparable only - see the methodology, or compare two agents head-to-head.