Electricity Bench

claude-code + claude-fable-5-1

claude-code harness · Claude Max 20x ($200/mo)

BOVERALL

Clears tier 4 of 5 (partial at tier 5) on Real-world issues

Capability mean across 3 graded suite families.

capability
73%
plan cost per run
28.0% of Claude Max 20x
median task time
1m 46s
reliability
100%
12345
What the tiers mean

"Clears" marks the highest rung where every task at it and below was solved. Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some. The headline stops at the last unbroken rung, so a solved tier above a half-solved one counts only as partial.

Clears level 3 of 5 (partial at levels 4 and 5) on Vibe coding v1 - the vaguest report level at which it still fixed every bug, having also cleared every richer level below
12345
What the levels mean

Level 1: a full brief naming the surface and the acceptance · level 3: the report as filed · level 5: a vague vibe report in a non-technical user's words. Context falls as the level rises, so vaguer is harder. A cleared level fixed every bug at it and below; a partial level fixed some. The headline stops at the last unbroken level, so a fix above a half-cleared level counts only as partial.

13 of this suite's tasks were credited rather than run: on a suite that asks the same problem at several levels of detail, solving it from the vaguest description credits the more detailed ones instead of asking again.

Per-suite results

C
Real-world issues v3
73%
plan cost per run
28.0% of Claude Max 20x
median task time
1m 46s
reliability
100%
tier 1tier 2tier 3tier 4tier 5

13 of 15 tasks solved - grouped by tier, easiest first; hover a square for its time and turns.

Run on claude-code 2.1.251 (Claude Code), model claude-fable-5-1 (--effort medium), week 2026-W36.

One run of this suite consumes 28.0% of Claude Max 20x's 5h window and 5.0% of the weekly limit.

Claude Max 20x's weekly limit holds about 7.2 fully spent 5h windows, out of the 33.6 a calendar week contains. Pooled over 618 runs in 21 session windows.

As work per window: about 522 suite runs per month on this plan - 2.61 runs per plan-dollar (at $200/mo, priced 2026-08-05). Work figures compare outputs and are comparable across agents; the raw quota share above is not.

How the plan cost was measured

Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Claude Max 20x account on eb-runner, shared with the operator IDE panes; quiet-window until a dedicated benchmark login exists.

C
Spec planning v1
74%
plan cost per run
15.0% of Claude Max 20x
median task time
6m 51s
reliability
100%

0 of 4 tasks solved - hover a square for its time and turns.

Run on claude-code 2.1.251 (Claude Code), model claude-fable-5-1 (--effort medium), week 2026-W36.

One run of this suite consumes 15.0% of Claude Max 20x's 5h window and 2.0% of the weekly limit.

Claude Max 20x's weekly limit holds about 7.2 fully spent 5h windows, out of the 33.6 a calendar week contains. Pooled over 618 runs in 21 session windows.

As work per window: about 974 suite runs per month on this plan - 4.87 runs per plan-dollar (at $200/mo, priced 2026-08-05). Work figures compare outputs and are comparable across agents; the raw quota share above is not.

How the plan cost was measured

Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Claude Max 20x account on eb-runner, shared with the operator IDE panes; quiet-window until a dedicated benchmark login exists.

B
Vibe coding v1
85%
plan cost per solve
1.8% of Claude Max 20x
median task time
3m 48s
reliability
100%
level 5level 4level 3level 2level 1issue aissue bissue cissue dissue e

24 of 25 tasks solved (13 credited without running) - each row is one bug, vaguest report first and context growing to the right; hover a square for its time and turns.

Run on claude-code 2.1.251 (Claude Code), model claude-fable-5-1 (--effort medium), week 2026-W36.

One good solve on this suite consumes, on average, 1.8% of Claude Max 20x's 5h window and 0.2% of the weekly limit.

Claude Max 20x's weekly limit holds about 7.2 fully spent 5h windows, out of the 33.6 a calendar week contains. Pooled over 618 runs in 21 session windows.

As work per window: about 8117 good solves per month on this plan - 40.58 solves per plan-dollar (at $200/mo, priced 2026-08-05). Work figures compare outputs and are comparable across agents; the raw quota share above is not.

How the plan cost was measured

Counted over each issue's first solve (5 solves); the walk's failed harder cells and confirmation runs are the benchmark's own search cost and are excluded (full walk: 21.0% over 12 cells). Measured per cell run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Claude Max 20x account on eb-runner, shared with the operator IDE panes; quiet-window until a dedicated benchmark login exists.

Task outputs from these runs are withheld: publishing them would reveal the private suite. Per-task scores, times and turns are in the chips above.

Trajectory

One point per grading week - a re-run within a week shows only its latest execution. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. Click a point or week for plan usage. All agents over time →

Suite
Real-world issues · 5 weeks graded · latest 2026-W36
50%60%70%80%2026-W322026-W332026-W342026-W352026-W36New registration - model claude-fable-5 -> claude-fable-5-1, same harness, plan and settingsclaude-fable-5 · 2026-W32 · Aug 5, 2026 · 60% · Cclaude-fable-5 · 2026-W33 · Aug 10, 2026 · 60% · Cclaude-fable-5 · 2026-W34 · Aug 17, 2026 · 60% · Cclaude-fable-5 · 2026-W35 · Aug 23, 2026 · 73% · Cclaude-code + claude-fable-5-1 · 2026-W36 · Sep 1, 2026 · 73% · C

Dotted = a new model version took the line over (grok-4.5 → grok-4.6): the points either side were graded by different models. Plan or settings re-registrations join solid.

Weekly heads

Newest week first. Click a row for usage.

WeekGradeCapabilityClearsTask timeUsage /1w
2026-W36 · v3nowC73%tier 4 · p51m 46s5.0% of Claude Max 20x
2026-W35 · v3C73%tier 4 · p53m 29s4.0% of Claude Max 20x
2026-W34 · v3C60%tier 4 · p53m 12s5.0% of Claude Max 20x
2026-W33 · v3C60%tier 4 · p53m 5s4.0% of Claude Max 20x
2026-W32 · v3C60%tier 4 · p52m 49s8.0% of Claude Max 5x

Frequently asked questions

Is Claude Fable 5.1 worth it on Claude Max 20x?

On our private real-world suite (real-world@v3), Claude Fable 5.1 earned grade C with 73% capability, clearing every problem up to tier 4 of 5 - a cross-cutting problem needing broad codebase orientation and landing some tier 5 problems. It runs on Claude Max 20x at $200/month (public price as of 2026-08-05). One full suite run consumed 28.0% of the plan's 5h window. Measured 2026-09-01.

How good is Claude Fable 5.1 at real-world coding?

Claude Fable 5.1 clears tier 4 of our five-tier real-world ladder - every task at that tier and below - where tier 1 is a localized single-file bug and tier 5 is a problem that took a human engineer hours. It solves some but not all tasks up at tier 5. Its capability score on real-world@v3 is 73% (grade C). Last benchmarked 2026-09-01.

How much coding work does Claude Max 20x buy with Claude Fable 5.1?

Roughly 522 suite runs per month on Claude Max 20x, or 2.61 runs per plan-dollar at $200/month (priced 2026-08-05). Usage is read from the provider's own meter per suite run; the raw quota share is an absolute figure on this plan, not comparable with another agent's plan.

Compare