Electricity Bench

2026-W39: Vibe coding

Week 39 of 2026 graded 13 coding agents on vibe coding. Claude Opus 5 leads with C+. The biggest move since W38 2026 is GPT-5.6 Luna, C- to F. Grades compare only within the same suite version, and plan usage is each agent's share of its own subscription. Every agent ran the same suite cut headless through its own tools.Measured

vibe-coding@v1 - the suite key these grades compare within

Published suite snapshot for 13 graded agents, current through Sep 22, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Vibe coding methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
Claude Opus 5claude-codeC+71.6%level 2 · partial level 52m 51s100%1.6%/5hClaude Max 5x
GPT-6 Astracodex-cliC-61.3%level 2 · partial level 52m 46s100%0.2%/1wChatGPT Pro 5x
Grok 4.6grok-cliC-61.3%level 2 · partial level 511m 4s100%0%/1wSuperGrok Plus
Kimi K3 256Kkimi-cliC-60.0%none6m 33s100%14%/5hKimi Allegretto
Claude Fable 5.1claude-codeC-58.7%level 3 · partial level 53m 58s100%6.2%/5hClaude Max 5x
Gemini 3.8 Flashantigravity-cliD+52.3%level 1 · partial level 56m 42s100%1.2%/1wGoogle AI Pro
Claude Sonnet 5claude-codeD+51.0%level 1 · partial level 54m 26s100%2.4%/5hClaude Max 5x
GPT-5.6 Solcodex-cliD+51.0%level 2 · partial level 53m 46s100%0.4%/1wChatGPT Pro 5x
Gemini 3.1 Proantigravity-cliF34.8%none4m 3s100%1.0%/1wGoogle AI Pro
GPT-5.6 Terracodex-cliF34.2%level 1 · partial level 55m 32s100%0%/1wChatGPT Pro 5x
Composer 2.5cursor-cliF31.6%none1m 56s100%0.2%/1mCursor Pro
GPT-5.6 Lunacodex-cliF25.2%level 1 · partial level 53m 40s100%0%/1wChatGPT Pro 5x
Claude Haiku 4.5claude-codeF22.6%none3m 36s100%2%/5hClaude Max 5x

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.