"Clears" marks the highest rung where every task at it and below was solved. Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some. The headline stops at the last unbroken rung, so a solved tier above a half-solved one counts only as partial.
Clears level 3 of 5 (partial at levels 4 and 5) on Vibe coding v1 - the vaguest report level at which it still fixed every bug, having also cleared every richer level below
12345
What the levels mean
Level 1: a full brief naming the surface and the acceptance · level 3: the report as filed · level 5: a vague vibe report in a non-technical user's words. Context falls as the level rises, so vaguer is harder. A cleared level fixed every bug at it and below; a partial level fixed some. The headline stops at the last unbroken level, so a fix above a half-cleared level counts only as partial.
13 of this suite's tasks were credited rather than run: on a suite that asks the same problem at several levels of detail, solving it from the vaguest description credits the more detailed ones instead of asking again.
13 of 15 tasks solved - grouped by tier, easiest first; hover a square for its time and turns. solved partial not solved
Run on claude-code 2.1.251 (Claude Code), model claude-fable-5-1 (--effort medium), week 2026-W36.
One run of this suite consumes 28.0% of Claude Max 20x's 5h window and 5.0% of the weekly limit.
Claude Max 20x's weekly limit holds about 7.2 fully spent 5h windows, out of the 33.6 a calendar week contains. Pooled over 618 runs in 21 session windows.
As work per window: about 522 suite runs per month on this plan - 2.61 runs per plan-dollar (at $200/mo, priced 2026-08-05). Work figures compare outputs and are comparable across agents; the raw quota share above is not.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Claude Max 20x account on eb-runner, shared with the operator IDE panes; quiet-window until a dedicated benchmark login exists.
0 of 4 tasks solved - hover a square for its time and turns. solved partial not solved
Run on claude-code 2.1.251 (Claude Code), model claude-fable-5-1 (--effort medium), week 2026-W36.
One run of this suite consumes 15.0% of Claude Max 20x's 5h window and 2.0% of the weekly limit.
Claude Max 20x's weekly limit holds about 7.2 fully spent 5h windows, out of the 33.6 a calendar week contains. Pooled over 618 runs in 21 session windows.
As work per window: about 974 suite runs per month on this plan - 4.87 runs per plan-dollar (at $200/mo, priced 2026-08-05). Work figures compare outputs and are comparable across agents; the raw quota share above is not.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Claude Max 20x account on eb-runner, shared with the operator IDE panes; quiet-window until a dedicated benchmark login exists.
24 of 25 tasks solved (13 credited without running) - each row is one bug, vaguest report first and context growing to the right; hover a square for its time and turns. solved partial not solved credited
Run on claude-code 2.1.251 (Claude Code), model claude-fable-5-1 (--effort medium), week 2026-W36.
One good solve on this suite consumes, on average, 1.8% of Claude Max 20x's 5h window and 0.2% of the weekly limit.
Claude Max 20x's weekly limit holds about 7.2 fully spent 5h windows, out of the 33.6 a calendar week contains. Pooled over 618 runs in 21 session windows.
As work per window: about 8117 good solves per month on this plan - 40.58 solves per plan-dollar (at $200/mo, priced 2026-08-05). Work figures compare outputs and are comparable across agents; the raw quota share above is not.
How the plan cost was measured
Counted over each issue's first solve (5 solves); the walk's failed harder cells and confirmation runs are the benchmark's own search cost and are excluded (full walk: 21.0% over 12 cells). Measured per cell run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Claude Max 20x account on eb-runner, shared with the operator IDE panes; quiet-window until a dedicated benchmark login exists.
Task outputs from these runs are withheld: publishing them would reveal the private suite. Per-task scores, times and turns are in the chips above.
Trajectory
One point per grading week - a re-run within a week shows only its latest execution. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. Click a point or week for plan usage. All agents over time →
Dotted = a new model version took the line over (grok-4.5 → grok-4.6): the points either side were graded by different models. Plan or settings re-registrations join solid.
Weekly heads
Newest week first. Click a row for usage.
Week
Grade
Capability
Clears
Task time
Usage /1w
2026-W36 · v3now
C
73%
tier 4 · p5
1m 46s
5.0% of Claude Max 20x
2026-W35 · v3
C
73%
tier 4 · p5
3m 29s
4.0% of Claude Max 20x
2026-W34 · v3
C
60%
tier 4 · p5
3m 12s
5.0% of Claude Max 20x
2026-W33 · v3
C
60%
tier 4 · p5
3m 5s
4.0% of Claude Max 20x
2026-W32 · v3
C
60%
tier 4 · p5
2m 49s
8.0% of Claude Max 5x
Frequently asked questions
Is Claude Fable 5.1 worth it on Claude Max 20x?
On our private real-world suite (real-world@v3), Claude Fable 5.1 earned grade C with 73% capability, clearing every problem up to tier 4 of 5 - a cross-cutting problem needing broad codebase orientation and landing some tier 5 problems. It runs on Claude Max 20x at $200/month (public price as of 2026-08-05). One full suite run consumed 28.0% of the plan's 5h window. Measured 2026-09-01.
How good is Claude Fable 5.1 at real-world coding?
Claude Fable 5.1 clears tier 4 of our five-tier real-world ladder - every task at that tier and below - where tier 1 is a localized single-file bug and tier 5 is a problem that took a human engineer hours. It solves some but not all tasks up at tier 5. Its capability score on real-world@v3 is 73% (grade C). Last benchmarked 2026-09-01.
How much coding work does Claude Max 20x buy with Claude Fable 5.1?
Roughly 522 suite runs per month on Claude Max 20x, or 2.61 runs per plan-dollar at $200/month (priced 2026-08-05). Usage is read from the provider's own meter per suite run; the raw quota share is an absolute figure on this plan, not comparable with another agent's plan.