Vibe coding
vibe-coding@v1 - the suite key used in the data and the URL
The context grid (electricitybench-226). Five real multi-fix bugs on forks of our own repositories, each asked at all five context levels - tier 1 a full detailed brief that names the surface and the acceptance, tier 5 a vague vibe report in a non-technical user's words - 25 tasks, tier = context level of the prompt, so the context axis is finally independent of which issue is being asked. Same fork-of-a-real-repo-at-the-pre-fix-commit format as real-world@v3; each issue's checker is the repo's own tests plus hidden assertions authored at factoring-neutral pre-fix boundaries, reused verbatim across its five cells. Substrate facts a fork cannot hold (what a live tool paints or leaves on disk) ride along at every level; the ladder thins diagnosis, not substrate. Tier weights 1/2/4/8/16 make the gradient and the headline is the derived "solves down to context level N" rung.
Context ladder
Every task sits on a difficulty tier, every subject runs every tier, and tier weights make the gradient: the capability score is dominated by the highest tiers cleared. The headline is the derived "clears tier N" - the highest rung where every task at it and below was solved, with a partial marker for the highest rung above it that the subject still landed a task on.
| Tier | Tasks | Weight per task | What sits here |
|---|---|---|---|
| 1 | 5 | 1 | full brief - names the surface and the acceptance criteria |
| 2 | 5 | 2 | detailed report from a technical user |
| 3 | 5 | 4 | the report as filed |
| 4 | 5 | 8 | thin report - a faint trace, no diagnosis |
| 5 | 5 | 16 | vibe report - a vague symptom in a non-technical user's words |
Headroom is designed, not broken. A top-tier task that no current agent solves is kept as long as its reference solution passes calibration: it is reachable and proven solvable, and its emptiness is the finding.
Grade bands
| Grade | Capability ≥ |
|---|---|
| A+ | 96.6667% |
| A | 93.3333% |
| A- | 90% |
| B+ | 85% |
| B | 80% |
| B- | 75% |
| C+ | 68.3333% |
| C | 61.6667% |
| C- | 55% |
| D+ | 48.3333% |
| D | 41.6667% |
| D- | 35% |
| F | below the lowest passing band |
This suite is marked outputs-reveal-prompts: its tasks publish scores and metadata only, because the outputs would otherwise reconstruct the prompts.
Task privacy
Every task in this suite stays private - publishing task content would let it leak into training data and grade memorization instead of capability. What is public is the method: the ladder shape, weights, scoring, and grade bands above, plus every graded run's scores and operational metadata.