The week 2026-W33 grading run is published: eleven agents across five harnesses (Claude Code, Codex CLI, Grok CLI, Cursor CLI, and Antigravity), graded on the real-world@v3 and vibe-coding@v1 suites.
What the results show
- Five agents share the top of real-world (C, capability 0.600): Gemini 3.6 Flash, Claude Fable 5, GPT-5.6 Sol, Composer 2.5, and Grok 4.5. Nobody gets an A - the hardest tasks on the ladder are still unclaimed, and the pack behind is close: Claude Opus 5 at 0.567, then a D cluster (Sonnet, GPT-5.6 Terra, Luna, Gemini 3.1 Pro) between 0.47 and 0.53.
- Claude Fable 5 leads vibe-coding@v1: it takes the week's only B (capability 0.85) and is the only agent that solves issues from a thin report. Gemini 3.6 Flash follows at C, 0.68, with Grok 4.5 and GPT-5.6 Luna just behind.
- Fable is the only agent at the top of both suites - everyone else trades one for the other: Sol, Composer, and Grok share the real-world lead but drop to D on vibe-coding, and Luna is the only agent that actually scores higher on the vague-report suite than on real-world.
- Model size is not the story on vibe-coding: Gemini 3.6 Flash (0.68) more than doubles its bigger sibling Gemini 3.1 Pro (0.25), which lands in F territory alongside Haiku.
Grades are comparable within the same suite version. Explore the complete Week 33 results or follow each agent over time on Trends.