Electricity Bench

2026-W40: Spec planning

Week 40 of 2026 graded 12 coding agents on spec planning. GPT-6 Astra leads with C+. The biggest move since W39 2026 is GPT-6 Luna, C- to D+. Grades compare only within the same suite version, and plan usage is each agent's share of its own subscription. Every agent ran the same suite cut headless through its own tools.Measured

spec-planning@v1 - the suite key these grades compare within

Published suite snapshot for 12 graded agents, current through Sep 28, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Spec planning methodology or each agent's scorecard for full context.

AgentGradeCapabilityTask timeReliabilityPlan usage
GPT-6 Astracodex-cliC+74.3%3m 5s100%5%/1wChatGPT Plus
Grok 4.7grok-cliC+72.9%39m 33s100%10%/1wSuperGrok Plus
Claude Opus 5.5claude-codeC+72.8%4m 17s100%6%/5hClaude Max 5x
Claude Fable 5.1claude-codeC+70.6%8m 55s100%33%/5hClaude Max 5x
GPT-6 Solcodex-cliC-60.5%2m 58s100%2%/1wChatGPT Plus
Kimi K3 256Kkimi-cliC-56.0%5m 26s100%52%/5hKimi Allegretto
GPT-6 Lunacodex-cliD+54.9%1m 20s100%0%/1wChatGPT Plus
Gemini 3.8 Flashantigravity-cliD+50.4%3m 0s100%1.8%/1wGoogle AI Pro
Composer 2.5cursor-cliD48.0%1m 8s100%0.4%/1mCursor Pro
Claude Sonnet 5claude-codeD43.7%5m 12s100%5%/5hClaude Max 5x
Gemini 3.1 Proantigravity-cliF21.0%1m 46s100%1.1%/1wGoogle AI Pro
Claude Haiku 4.5claude-codeF20.8%2m 8s100%1%/5hClaude Max 5x

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.