2026-W33: Real-world issues
real-world@v3 - the suite key these grades compare within
Published suite snapshot for 11 graded agents, current through Aug 11, 2026.
Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Real-world issues methodology or each agent's scorecard for full context.
| Agent | Grade | Capability | Clears | Task time | Reliability | Plan usage |
|---|---|---|---|---|---|---|
| claude-fable-5claude-code | C | 60.0% | tier 4 · partial tier 5 | 3m 5s | 100% | 34%/5hClaude Max 20x |
| gpt-5.6-solcodex-cli | C | 60.0% | tier 4 · partial tier 5 | 2m 57s | 100% | 2.8%/1wChatGPT Plus |
| composer-2.5cursor-cli | C | 60.0% | tier 4 · partial tier 5 | 3m 52s | 100% | 1.9%/1mCursor Pro |
| grok-4.5grok-cli | C | 60.0% | tier 4 · partial tier 5 | 3m 9s | 100% | 6.9%/1wSuperGrok |
| claude-opus-5claude-code | C | 56.7% | tier 2 · partial tier 5 | 3m 57s | 100% | 11%/5hClaude Max 20x |
| claude-sonnet-5claude-code | D | 53.3% | tier 3 · partial tier 5 | 5m 6s | 100% | 9.7%/5hClaude Max 20x |
| gpt-5.6-terracodex-cli | D | 53.3% | tier 3 · partial tier 5 | 1m 57s | 100% | 2.1%/1wChatGPT Plus |
| gemini-3.6-flashantigravity-cli | D | 52.0% | tier 4 · partial tier 5 | 2m 19s | 100% | 10.3%/1wGoogle AI Pro |
| gpt-5.6-lunacodex-cli | D | 50.8% | none | 2m 15s | 100% | 3.0%/1wChatGPT Plus |
| gemini-3.1-proantigravity-cli | D | 46.7% | tier 3 · partial tier 5 | 4m 24s | 100% | 8.0%/1wGoogle AI Pro |
| claude-haiku-4-5claude-code | F | 28.3% | tier 1 · partial tier 5 | 3m 31s | 100% | 3%/5hClaude Max 20x |
This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.