The leaderboard now shows what it was built to show: real agent harnesses, on the subscriptions people actually pay for, graded on the same private task ladder - with the quota each run burned read from the provider's own meter.
This first weekly run graded ten agents across five harnesses (Claude Code, Codex CLI, Grok CLI, Cursor CLI, and Antigravity) on the real-world@v3 suite: fifteen private tasks forked from real repositories at the commit before a real fix landed, arranged on a five-tier difficulty ladder from localized single-file bugs to problems that took a human maintainer hours.
A few honest headlines from the run:
- Four agents share the top grade (C, capability 0.600): Claude Opus 5 and Claude Fable 5 on Claude Max 5x, Grok 4.5 on SuperGrok, and - the surprise of the run - Gemini 3.6 Flash on Google AI Pro. Nobody gets an A: the top tier holds tasks no current agent clears, and that headroom is deliberate.
- Price does not predict rank. A $20 Google AI Pro flash model and a $30 SuperGrok plan sit level with a $100 Max plan, while some famous names land lower. The scorecards show exactly which rungs each agent cleared.
- Quota is the second axis. The same suite pass costs anywhere from ~3% of a ChatGPT Plus week to ~64% of a Claude Max 5-hour window to ~20% of a Google AI Pro week - all read from each provider's own meter, never estimated from marketing numbers.
Every grade links its full scorecard: per-tier ladder results, latency, reliability, the exact harness version pinned, and the plan-quota consumption with its measurement method disclosed. Where a reading could not be measured, the card says so instead of showing a number we cannot stand behind.
The cadence from here is weekly: every Monday the harnesses upgrade to their latest releases and every subject re-grades, so the board always reflects the product you would install today - and the trends page grows a trajectory point per agent per week.
If you want the weekly readout in your inbox, the newsletter now has working double-opt-in.