Private suites · Public grades
What a coding‑agent subscription is actually worth
We run every coding agent on private task suites, on the plan you would actually buy. Two numbers per agent: how hard a problem it can take, and what a task costs in the plan's own quota.
Choosing a plan? How to choose · Two-minute quiz
Latest
All news- GPT-6.1 Sol releaseGPT-6.1 Sol is good, just two weeks too late
- Claude Sonnet 5.5 releaseClaude Sonnet 5.5 is a Sonnet done right
- W40 grading runClaude pulls further ahead as GPT-6 agents climb - Week 40
LeaderboardW40 2026
Last updated
| # | Agent | Overall | Plan usage | Approx. cost | Real-world issues | Spec planning | Vibe coding |
|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5.5claude-code | B- | 2% of weekly limitClaude Max 5x2% of weekly limitClaude Max 5x~10% of weekly limitestClaude Pro2% of weekly limitClaude Max 5x~1% of weekly limitestClaude Max 20x | $0.024/task$0.024/task≈$0.024/taskest$0.024/task≈$0.024/taskest | C+ | C+ | B+ |
| 2 | Claude Fable 5.1claude-code | B- | 8% of weekly limitClaude Max 5x8% of weekly limitClaude Max 5xnot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.8% of weekly limitClaude Max 5x~4% of weekly limitestClaude Max 20x | $0.097/task$0.097/tasknot availableAnthropic does not offer Fable on Claude Pro - it ships only on the Max tiers, so there is no $20 plan to run this agent on and nothing to estimate.$0.097/task≈$0.097/taskest | B+ | C+ | C+ |
| 3 | GPT-6.1 Solcodex-cli | C | ~1.2% of weekly limitestChatGPT Pro 5x6% of weekly limitChatGPT Plus6% of weekly limitChatGPT Plus~1.2% of weekly limitestChatGPT Pro 5x~0.3% of weekly limitestChatGPT Pro 20xnot available | ≈$0.015/taskest$0.015/task$0.015/task≈$0.015/taskest≈$0.007/taskest | C- | C+ | C+ |
| 4 | Grok 4.7grok-cli | C | 15% of weekly limitSuperGrok Plus15% of weekly limitSuperGrok Plusno estimate15% of weekly limitSuperGrok Plusno estimate | $0.182/task$0.182/task-$0.182/task- | C- | C+ | C |
| 5 | GPT-6 Astracodex-cli | C | ~4.4% of weekly limitestChatGPT Pro 5x22% of weekly limitChatGPT Plus22% of weekly limitChatGPT Plus~4.4% of weekly limitestChatGPT Pro 5x~1.1% of weekly limitestChatGPT Pro 20xnot available | ≈$0.053/taskest$0.053/task$0.053/task≈$0.053/taskest≈$0.027/taskest | D | C+ | C+ |
| 6 | Claude Sonnet 5.5claude-code | C- | <1% of weekly limitClaude Max 5x<1% of weekly limitClaude Max 5x~<5% of weekly limitestClaude Pro<1% of weekly limitClaude Max 5x~<0.5% of weekly limitestClaude Max 20x | $0.003/task$0.003/task≈$0.003/taskest$0.003/task≈$0.003/taskest | C- | C+ | D+ |
| 7 | Gemini 3.8 Flashantigravity-cli | C- | 23.7% of weekly limitGoogle AI Pro23.7% of weekly limitGoogle AI Pro23.7% of weekly limitGoogle AI Pro~4.7% of weekly limitestGoogle AI Ultra 5x~1.2% of weekly limitestGoogle AI Ultra 20x | $0.057/task$0.057/task$0.057/task≈$0.057/taskest≈$0.029/taskest | C- | D+ | C- |
| 8 | Kimi K3 256Kkimi-cli | D+ | 31% of weekly limitKimi Allegretto31% of weekly limitKimi Allegrettono estimateno estimateno estimate | $0.146/task$0.146/task--- | C- | C- | F |
| 9 | GPT-6 Lunacodex-cli | D | ~<0.2% of weekly limitestChatGPT Pro 5x<1% of weekly limitChatGPT Plus<1% of weekly limitChatGPT Plus~<0.2% of weekly limitestChatGPT Pro 5x~<0.1% of weekly limitestChatGPT Pro 20xnot available | ----- | C- | D+ | F |
| 10 | GPT-5.6 Terracodex-cli | D | ~1.6% of weekly limitestChatGPT Pro 5x8% of weekly limitChatGPT Plus8% of weekly limitChatGPT Plus~1.6% of weekly limitestChatGPT Pro 5x~0.4% of weekly limitestChatGPT Pro 20xnot available | ≈$0.019/taskest$0.019/task$0.019/task≈$0.019/taskest≈$0.010/taskest | D+ | D+ | F |
| 11 | Composer 2.5cursor-cli | D | 2.0% of monthly limitCursor Pro2.0% of monthly limitCursor Pro2.0% of monthly limitCursor Pro~0.7% of monthly limitestCursor Pro+~0.1% of monthly limitestCursor Ultra | $0.021/task$0.021/task$0.021/task≈$0.021/taskest≈$0.011/taskest | C- | D | F |
| 12 | Gemini 3.1 Proantigravity-cli | F | 15.9% of weekly limitGoogle AI Pro15.9% of weekly limitGoogle AI Pro15.9% of weekly limitGoogle AI Pro~3.2% of weekly limitestGoogle AI Ultra 5x~0.8% of weekly limitestGoogle AI Ultra 20x | $0.038/task$0.038/task$0.038/task≈$0.038/taskest≈$0.019/taskest | D+ | F | F |
| 13 | Claude Haiku 4.5claude-code | F | 1% of weekly limitClaude Max 5x1% of weekly limitClaude Max 5x~5% of weekly limitestClaude Pro1% of weekly limitClaude Max 5x~0.5% of weekly limitestClaude Max 20x | $0.012/task$0.012/task≈$0.012/taskest$0.012/task≈$0.012/taskest | D- | F | F |
Default view: Claude rows are estimated at Claude Max 5x (~$100) from the Max 20x run they were measured on, marked est. The weekly limit grows ~2x from Max 5x to Max 20x, not the advertised session 4x. OpenAI rows are measured on ChatGPT Pro 5x. Every other row reads on the plan it ran on. Pick "Measured plan" for every row as measured, or a budget band to re-express all of them.
Budget view: rows whose measured plan sits in the picked band keep their measured figure; the rest are scaled from the measured plan by the provider's own advertised tier multiples (marked est) - approximate, never measured. Providers that publish no multiple get no estimate. Our paired Claude Max 5x/20x runs measured ~5.7x where 4x is advertised, so treat estimates as rough. A model the tier does not sell at all - Fable on Claude Pro - reads not available rather than an estimate.
Grades are within-suite-version comparable only. See methodology. Or compare two agents head-to-head. Overall is the capability mean across a subject's graded suite families.
Popular comparisons
B-vsC
Claude Fable 5.1 vs Kimi K3 256KB-vsD+
Claude Opus 5.5 vs Grok 4.7B-vsC
Gemini 3.8 Flash vs Claude Opus 5.5C-vsB-
Claude Fable 5.1 vs Claude Haiku 4.5B-vsF
Claude Fable 5.1 vs Claude Opus 5.5B-vsB-
Claude Max 5x vs ChatGPT Pro 5xSubscription
ChatGPT Plus vs Claude ProSubscription
How it works
Task suites stay private so they cannot leak into training data - what is public is the methodology: ladder shape, weights, scoring, and grade bands.
Each agent runs every suite headless through its own tools, with its harness version and plan pinned: capability, plan cost, task time, and reliability, once per task per weekly run.
Only the graded outcome is published. Third-party claims are labeled context, never blended into our grades.