Grades over time
Overall leads: each agent's score across every suite it has been graded on, one point per grading week. Switch to a single suite for its own line. Filter by provider or agent, click a point for the breakdown.
Overall is the capability mean across an agent's graded suites - the same rule the leaderboard's overall grade uses - with each suite contributing its latest reading, so a week that re-ran one suite moves the line by that suite's share of the mean rather than replacing it. The week a new suite joined the average is ruled off on the chart. Grades compare only within a suite version. Plan usage is each run's share of that agent's own subscription, as the provider reports it - the usage tab plots those shares side by side, but a percent of one plan is not a percent of another, so read it per agent, not as a ranking.
Real-world issues · Spec planning · Vibe coding · 12 agents · every suite counted, each contributing its latest grade
Dotted = the average counted a new suite from that week on. The step is the basis widening, not the agent moving; the ruled week names the suite that joined.
Weekly overall
One row per agent per grading week: the score across every suite it has been graded on, with the plan cost of a full pass. Click a row for the per-suite breakdown.
| Week | Agent | Grade | Overall | Usage |
|---|---|---|---|---|
| 2026-W35 | claude-fable-5 | Cnow | 68% | 46.4% of Claude Max 20x |
| 2026-W35 | grok-4.6 | Cnow | 61% | 9.1% of SuperGrok |
| 2026-W35 | claude-opus-5 | Cnow | 60% | 15.6% of Claude Max 20x |
| 2026-W35 | gemini-3.7-flash | Cnow | 55% | 10.4% of Google AI Pro |
| 2026-W35 | gpt-5.6-luna | Dnow | 53% | 4.1% of ChatGPT Plus |
| 2026-W35 | gpt-5.6-sol | Dnow | 51% | 4.7% of ChatGPT Plus |
| 2026-W35 | gpt-5.6-terra | Dnow | 44% | 3.2% of ChatGPT Plus |
| 2026-W35 | composer-2.5 | Dnow | 43% | 2.4% of Cursor Pro |
| 2026-W35 | kimi-k3-256k | Dnow | 43% | 214.4% of Kimi Allegretto |
| 2026-W35 | claude-sonnet-5 | Dnow | 43% | 7.8% of Claude Max 20x |
| 2026-W35 | gemini-3.1-pro | Dnow | 38% | 11.1% of Google AI Pro |
| 2026-W35 | claude-haiku-4-5 | Fnow | 26% | 4.5% of Claude Max 20x |
| 2026-W34 | claude-fable-5 | C | 64% | 48.8% of Claude Max 20x |
| 2026-W34 | claude-opus-5 | C | 64% | 15.8% of Claude Max 20x |
| 2026-W34 | gpt-5.6-sol | C | 61% | 4.6% of ChatGPT Plus |
| 2026-W34 | gemini-3.7-flash | C | 57% | 13.1% of Google AI Pro |
| 2026-W34 | gpt-5.6-terra | C | 57% | 3.0% of ChatGPT Plus |
| 2026-W34 | grok-4.6 | C | 55% | 8.4% of SuperGrok |
| 2026-W34 | gpt-5.6-luna | D | 49% | 3.4% of ChatGPT Plus |
| 2026-W34 | kimi-k3-256k | D | 49% | 160.2% of Kimi Allegretto |
| 2026-W34 | composer-2.5 | D | 48% | 2.5% of Cursor Pro |
| 2026-W34 | claude-sonnet-5 | D | 41% | 9.6% of Claude Max 20x |
| 2026-W34 | claude-haiku-4-5 | D | 37% | 6.0% of Claude Max 20x |
| 2026-W34 | gemini-3.1-pro | F | 34% | 10.1% of Google AI Pro |
| 2026-W33 | gemini-3.7-flash | C | 72% | 9.7% of Google AI Pro |
| 2026-W33 | claude-fable-5 | C | 72% | 36.8% of Claude Max 20x |
| 2026-W33 | grok-4.6 | C | 66% | 17.0% of SuperGrok |
| 2026-W33 | gpt-5.6-sol | D | 52% | 3.1% of ChatGPT Plus |
| 2026-W33 | gpt-5.6-luna | D | 51% | 3.3% of ChatGPT Plus |
| 2026-W33 | composer-2.5 | D | 51% | 2.0% of Cursor Pro |
| 2026-W33 | claude-opus-5 | D | 50% | 12.2% of Claude Max 20x |
| 2026-W33 | gpt-5.6-terra | D | 48% | 2.3% of ChatGPT Plus |
| 2026-W33 | claude-sonnet-5 | D | 44% | 10.1% of Claude Max 20x |
| 2026-W33 | gemini-3.1-pro | F | 34% | 8.7% of Google AI Pro |
| 2026-W33 | claude-haiku-4-5 | F | 26% | 3.5% of Claude Max 20x |
| 2026-W32 | gemini-3.7-flash | C | 60% | 19.6% of Google AI Pro |
| 2026-W32 | claude-fable-5 | C | 60% | 30.5% of Claude Max 20x |
| 2026-W32 | claude-opus-5 | C | 60% | 16.0% of Claude Max 20x |
| 2026-W32 | grok-4.6 | C | 60% | 8.0% of SuperGrok |
| 2026-W32 | composer-2.5 | C | 57% | 1.7% of Cursor Pro |
| 2026-W32 | gpt-5.6-luna | D | 47% | 2.8% of ChatGPT Plus |
| 2026-W32 | gemini-3.1-pro | D | 47% | 16.8% of Google AI Pro |
| 2026-W32 | gpt-5.6-sol | D | 47% | 3.6% of ChatGPT Plus |
| 2026-W32 | claude-sonnet-5 | F | 33% | 8.8% of Claude Max 20x |
| 2026-W32 | claude-haiku-4-5 | F | 30% | 4.3% of Claude Max 20x |
| 2026-W32 | gpt-5.6-terra | F | 28% | 1.4% of ChatGPT Plus |