2026-W34: Real-world issues
Week 34 of 2026 graded 12 coding agents on real-world issues. Claude Haiku 4.5 leads with C. The biggest move since W33 2026 is Claude Haiku 4.5, F to C. Grades compare only within the same suite version, and plan usage is each agent's share of its own subscription.Measured
real-world@v3 - the suite key these grades compare within
Published suite snapshot for 12 graded agents, current through Aug 18, 2026.
Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Real-world issues methodology or each agent's scorecard for full context.
| Agent | Grade | Capability | Clears | Task time | Reliability | Plan usage |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5claude-code | C | 63.3% | tier 2 · partial tier 5 | 3m 58s | 100% | 20%/5hClaude Max 5x |
| Claude Fable 5.1claude-code | C- | 60.0% | tier 4 · partial tier 5 | 3m 12s | 100% | 136%/5hClaude Max 5x |
| Claude Opus 5.5claude-code | C- | 60.0% | tier 4 · partial tier 5 | 3m 57s | 100% | 44%/5hClaude Max 5x |
| GPT-6.1 Solcodex-cli | C- | 60.0% | tier 4 · partial tier 5 | 2m 28s | 100% | 11.0%/1wChatGPT Plus |
| Grok 4.7grok-cli | C- | 60.0% | tier 4 · partial tier 5 | 6m 19s | 100% | 0.8%/1wSuperGrok Plus |
| Composer 2.5cursor-cli | C- | 56.7% | tier 2 · partial tier 5 | 1m 30s | 100% | 1.8%/1mCursor Pro |
| Kimi K3 256Kkimi-cli | C- | 56.7% | tier 2 · partial tier 5 | 3m 3s | 100% | 99%/5hKimi Allegretto |
| Gemini 3.1 Proantigravity-cli | D+ | 53.3% | tier 3 · partial tier 5 | 3m 25s | 100% | 8.5%/1wGoogle AI Pro |
| Gemini 3.8 Flashantigravity-cli | D+ | 53.3% | tier 3 · partial tier 5 | 2m 50s | 100% | 11.3%/1wGoogle AI Pro |
| GPT-5.6 Terracodex-cli | D+ | 53.3% | tier 3 · partial tier 5 | 1m 49s | 100% | 7.7%/1wChatGPT Plus |
| GPT-6 Lunacodex-cli | D+ | 53.3% | tier 3 · partial tier 5 | 1m 59s | 100% | 0.8%/1wChatGPT Plus |
| Claude Sonnet 5.5claude-code | D+ | 50.0% | tier 2 · partial tier 5 | 4m 22s | 100% | 24%/5hClaude Max 5x |
This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.