The week 2026-W36 grading run is published: no newcomers across real-world@v3, spec-planning@v1, and vibe-coding@v1.
What the results show
- Claude Fable 5 stays the agent to beat, and the whole Claude family got quicker - a median task that took about six minutes now finishes in under five.
- OpenAI added a 5-hour usage window to the $20 ChatGPT Plus plan. How much an agent can do in one sitting now matters as much as what a week's allowance buys, so we are adding that ratio to the bench for every agent whose plan sets limits over more than one window.
- GPT 5.6 Sol had its strongest week yet, up on real-world work and planning, and paid for it in time - each task takes it about a minute longer than last week. Grok 4.6 is still the slowest agent on the bench.
The standings barely move, but the Plus cap shifts the plan math a little. Our take: the $20 Claude plan gets more attractive wherever Sol is not needed, and anyone who works in one or two long sittings a week should no longer pick OpenAI's entry plan.
Grades are comparable within the same suite version. Explore the complete Week 36 results or follow each agent over time on Trends.