No new models this week. The week 2026-W38 grading run, graded September 13 to 14, covers the same 13 agents on real-world@v3, spec-planning@v1 and vibe-coding@v1.
What the results show
- The top belongs to Claude, with GPT-6 Astra closing the gap. If you keep one subscription for everything, Claude Fable 5.1 is still the safest pick, near the top of every kind of work even after missing a hard repository issue it solved last week. Claude Opus 5 took over real repository work outright, clearing issues it used to miss, though it got worse when you hand it a one-line bug report and let it run. We re-ran its suspect bug-report tasks and the misses are real. Alongside the two, GPT-6 Astra keeps gaining ground. It now writes the strongest plans together with Fable and leads the bug-report work, with real repository issues still its weak spot.
- Kimi K3 answers twice as fast. Same grades as last week at half the wait. A repository issue now takes Kimi K3 about a minute and a half at the median, down from over three. If the wait was your reason to skip it, that reason is fading. The allowance drain that keeps it off our picks is unchanged.
- There is still no cheap daily driver. Gemini 3.8 Flash slipped on real repository issues and one full pass over them still eats about a fifth of its Google AI Pro week. We keep reaching for 3.7 Flash on small tasks (our take).
Not much changed on the board this week. Claude's limits did change. The +50% weekly promotion ended on September 13 and the permanent limits that replaced it on September 14 sit about 17% below what promo subscribers had (change log). The grades and usage in this run are unaffected. Claude usage on the board is measured on the five-hour window and those limits stayed the same. What changed is the weekly capacity behind each plan, so we will run the Claude models again midweek to check plan value under the new limits.
Grades are comparable within the same suite version. Explore the complete Week 38 results or follow each agent over time on Trends.