Electricity Bench

Real-world issues

real-world@v3 - the suite key used in the data and the URL

The difficulty ladder, re-authored so the top rung separates a strong harness from grok-4.5 (electricitybench-171, after v2 saturated). Same fork-of-a-real-repo-at-the-pre-fix-commit format as v2; the task is the original symptom report and the checker is the real fix's tests plus hidden assertions; tier weights 1/2/4/8/16 make the gradient and the headline is the derived "solves up to tier N" rung.

scoring
programmatic
repeats
1

Difficulty ladder

Every task sits on a difficulty tier, every subject runs every tier, and tier weights make the gradient: the capability score is dominated by the highest tiers cleared. The headline is the derived "clears tier N" - the highest rung where every task at it and below was solved, with a partial marker for the highest rung above it that the subject still landed a task on.

TierTasksWeight per taskWhat sits here
121a localized single-file bug; the symptom names the surface it appears on
232a single-service bug, symptom one hop from the cause
324cause in a different module than the symptom; multi-file fix
438cross-boundary: a contract mismatch between services, a fix spanning conventions
5516problems that took a human hours: races, misleading symptoms, fixes needing a design choice

Headroom is designed, not broken. A top-tier task that no current agent solves is kept as long as its reference solution passes calibration: it is reachable and proven solvable, and its emptiness is the finding.

Grade bands

GradeCapability ≥
A+96.6667%
A93.3333%
A-90%
B+85%
B80%
B-75%
C+68.3333%
C61.6667%
C-55%
D+48.3333%
D41.6667%
D-35%
Fbelow the lowest passing band

This suite is marked outputs-reveal-prompts: its tasks publish scores and metadata only, because the outputs would otherwise reconstruct the prompts.

Task privacy

Every task in this suite stays private - publishing task content would let it leak into training data and grade memorization instead of capability. What is public is the method: the ladder shape, weights, scoring, and grade bands above, plus every graded run's scores and operational metadata.