2026-W35: Spec planning
spec-planning@v1 - the suite key these grades compare within
Published suite snapshot for 12 graded agents, current through Aug 24, 2026.
Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Spec planning methodology or each agent's scorecard for full context.
| Agent | Grade | Capability | Task time | Reliability | Plan usage |
|---|---|---|---|---|---|
| grok-4.6grok-cli | C | 67.4% | 35m 46s | 100% | 5.5%/1wSuperGrok |
| gpt-5.6-solcodex-cli | C | 66.1% | 5m 19s | 100% | 0.6%/1wChatGPT Plus |
| claude-fable-5claude-code | C | 62.8% | 5m 54s | 100% | 11%/5hClaude Max 20x |
| claude-opus-5claude-code | C | 62.8% | 6m 1s | 100% | 4%/5hClaude Max 20x |
| gpt-5.6-lunacodex-cli | C | 57.4% | 2m 28s | 100% | 0.4%/1wChatGPT Plus |
| gpt-5.6-terracodex-cli | C | 55.9% | 2m 27s | 100% | 0.5%/1wChatGPT Plus |
| gemini-3.7-flashantigravity-cli | D | 44.1% | 1m 56s | 100% | 1.3%/1wGoogle AI Pro |
| composer-2.5cursor-cli | D | 42.8% | 2m 35s | 100% | 0.5%/1mCursor Pro |
| claude-sonnet-5claude-code | D | 39.3% | 5m 34s | 100% | 2%/5hClaude Max 20x |
| kimi-k3-256kkimi-cli | D | 38.0% | 6m 12s | 100% | 56%/5hKimi Allegretto |
| claude-haiku-4-5claude-code | F | 22.6% | 2m 15s | 100% | 0%/5hClaude Max 20x |
| gemini-3.1-proantigravity-cli | F | 20.8% | 2m 35s | 100% | 1.0%/1wGoogle AI Pro |
This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.