Electricity Bench

2026-W35: Spec planning

spec-planning@v1 - the suite key these grades compare within

Published suite snapshot for 12 graded agents, current through Aug 24, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Spec planning methodology or each agent's scorecard for full context.

AgentGradeCapabilityTask timeReliabilityPlan usage
grok-4.6grok-cliC67.4%35m 46s100%5.5%/1wSuperGrok
gpt-5.6-solcodex-cliC66.1%5m 19s100%0.6%/1wChatGPT Plus
claude-fable-5claude-codeC62.8%5m 54s100%11%/5hClaude Max 20x
claude-opus-5claude-codeC62.8%6m 1s100%4%/5hClaude Max 20x
gpt-5.6-lunacodex-cliC57.4%2m 28s100%0.4%/1wChatGPT Plus
gpt-5.6-terracodex-cliC55.9%2m 27s100%0.5%/1wChatGPT Plus
gemini-3.7-flashantigravity-cliD44.1%1m 56s100%1.3%/1wGoogle AI Pro
composer-2.5cursor-cliD42.8%2m 35s100%0.5%/1mCursor Pro
claude-sonnet-5claude-codeD39.3%5m 34s100%2%/5hClaude Max 20x
kimi-k3-256kkimi-cliD38.0%6m 12s100%56%/5hKimi Allegretto
claude-haiku-4-5claude-codeF22.6%2m 15s100%0%/5hClaude Max 20x
gemini-3.1-proantigravity-cliF20.8%2m 35s100%1.0%/1wGoogle AI Pro

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.