real-world@v1
Issues we already fixed, replayed on forks of the real repos frozen at the pre-fix commit. The task is the original symptom report; the checker is the real fix's tests plus hidden assertions.
scoring
programmatic
repeats
1
Grade bands
| Grade | Capability ≥ |
|---|---|
| A | 0.9 |
| B | 0.75 |
| C | 0.55 |
| D | 0.35 |
| F | below D |
This suite is marked outputs-reveal-prompts: its non-sample tasks publish scores and metadata only, because the outputs would otherwise reconstruct the prompts.
Sample task
Published in full, exactly as the models see it. All other tasks in the suite stay private.
No sample task published.