Electricity Bench

Vibe coding

vibe-coding@v1 - the suite key used in the data and the URL

The context grid (electricitybench-226). Five real multi-fix bugs on forks of our own repositories, each asked at all five context levels - tier 1 a full detailed brief that names the surface and the acceptance, tier 5 a vague vibe report in a non-technical user's words - 25 tasks, tier = context level of the prompt, so the context axis is finally independent of which issue is being asked. Same fork-of-a-real-repo-at-the-pre-fix-commit format as real-world@v3; each issue's checker is the repo's own tests plus hidden assertions authored at factoring-neutral pre-fix boundaries, reused verbatim across its five cells. Substrate facts a fork cannot hold (what a live tool paints or leaves on disk) ride along at every level; the ladder thins diagnosis, not substrate. Tier weights 1/2/4/8/16 make the gradient and the headline is the derived "solves down to context level N" rung.

scoring
programmatic
repeats
1

Context ladder

Every task sits on a difficulty tier, every subject runs every tier, and tier weights make the gradient: the capability score is dominated by the highest tiers cleared. The headline is the derived "clears tier N" - the highest rung where every task at it and below was solved, with a partial marker for the highest rung above it that the subject still landed a task on.

TierTasksWeight per taskWhat sits here
151full brief - names the surface and the acceptance criteria
252detailed report from a technical user
354the report as filed
458thin report - a faint trace, no diagnosis
5516vibe report - a vague symptom in a non-technical user's words

Headroom is designed, not broken. A top-tier task that no current agent solves is kept as long as its reference solution passes calibration: it is reachable and proven solvable, and its emptiness is the finding.

Grade bands

GradeCapability ≥
A0.9
B0.75
C0.55
D0.35
Fbelow D

This suite is marked outputs-reveal-prompts: its tasks publish scores and metadata only, because the outputs would otherwise reconstruct the prompts.

Task privacy

Every task in this suite stays private - publishing task content would let it leak into training data and grade memorization instead of capability. What is public is the method: the ladder shape, weights, scoring, and grade bands above, plus every graded run's scores and operational metadata.