Electricity Bench

2026-W40: Vibe coding

Week 40 of 2026 graded 12 coding agents on vibe coding. Claude Opus 5.5 leads with B+. The biggest move since W39 2026 is Claude Sonnet 5, D+ to F. Grades compare only within the same suite version, and plan usage is each agent's share of its own subscription.Measured

vibe-coding@v1 - the suite key these grades compare within

Published suite snapshot for 12 graded agents, current through Sep 28, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Vibe coding methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
Claude Opus 5.5claude-codeB+89.7%level 4 · partial level 52m 55s100%1%/5hClaude Max 5x
Claude Fable 5.1claude-codeC+74.2%level 3 · partial level 54m 29s100%4.2%/5hClaude Max 5x
GPT-6 Astracodex-cliC+74.2%level 3 · partial level 52m 40s100%1%/1wChatGPT Plus
Grok 4.7grok-cliC63.9%level 3 · partial level 511m 34s100%0.6%/1wSuperGrok Plus
GPT-6 Solcodex-cliC-61.3%level 2 · partial level 52m 41s100%0.4%/1wChatGPT Plus
Kimi K3 256Kkimi-cliC-58.7%level 3 · partial level 57m 41s100%13.2%/5hKimi Allegretto
Gemini 3.8 Flashantigravity-cliC-56.1%level 2 · partial level 57m 58s100%1.1%/1wGoogle AI Pro
Claude Sonnet 5claude-codeF31.0%none4m 31s100%1%/5hClaude Max 5x
GPT-6 Lunacodex-cliF31.0%none1m 24s100%0%/1wChatGPT Plus
Gemini 3.1 Proantigravity-cliF27.1%none4m 50s100%0.7%/1wGoogle AI Pro
Composer 2.5cursor-cliF22.6%level 1 · partial level 51m 37s100%0.1%/1mCursor Pro
Claude Haiku 4.5claude-codeF21.9%none4m 4s100%1.3%/5hClaude Max 5x

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.