Vibe coding
vibe-coding@v1 - the suite key used in the data and the URL
The context grid (electricitybench-226). Five real multi-fix bugs on forks of our own repositories, each asked at all five context levels - tier 1 a full detailed brief that names the surface and the acceptance, tier 5 a vague vibe report in a non-technical user's words - 25 tasks, tier = context level of the prompt, so the context axis is finally independent of which issue is being asked. Same fork-of-a-real-repo-at-the-pre-fix-commit format as real-world@v3; each issue's checker is the repo's own tests plus hidden assertions authored at factoring-neutral pre-fix boundaries, reused verbatim across its five cells. Substrate facts a fork cannot hold (what a live tool paints or leaves on disk) ride along at every level; the ladder thins diagnosis, not substrate. Tier weights 1/2/4/8/16 make the gradient and the headline is the derived "solves down to context level N" rung.
Context ladder
Every task sits on a difficulty tier, every subject runs every tier, and tier weights make the gradient: the capability score is dominated by the highest tiers cleared. The headline is the derived "clears tier N" - the highest rung where every task at it and below was solved, with a partial marker for the highest rung above it that the subject still landed a task on.
| Tier | Tasks | Weight per task | What sits here |
|---|---|---|---|
| 1 | 5 | 1 | full brief - names the surface and the acceptance criteria |
| 2 | 5 | 2 | detailed report from a technical user |
| 3 | 5 | 4 | the report as filed |
| 4 | 5 | 8 | thin report - a faint trace, no diagnosis |
| 5 | 5 | 16 | vibe report - a vague symptom in a non-technical user's words |
Headroom is designed, not broken. A top-tier task that no current agent solves is kept as long as its reference solution passes calibration: it is reachable and proven solvable, and its emptiness is the finding.
Grade bands
| Grade | Capability ≥ |
|---|---|
| A | 0.9 |
| B | 0.75 |
| C | 0.55 |
| D | 0.35 |
| F | below D |
This suite is marked outputs-reveal-prompts: its tasks publish scores and metadata only, because the outputs would otherwise reconstruct the prompts.
Task privacy
Every task in this suite stays private - publishing task content would let it leak into training data and grade memorization instead of capability. What is public is the method: the ladder shape, weights, scoring, and grade bands above, plus every graded run's scores and operational metadata.