Electricity Bench

2026-W37: Vibe coding

vibe-coding@v1 - the suite key these grades compare within

Published suite snapshot for 13 graded agents, current through Sep 7, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Vibe coding methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
GPT-6 Astracodex-cliC74.2%level 3 · partial level 53m 0s100%0%/1wChatGPT Pro $100
Claude Fable 5.1claude-codeC69.0%level 3 · partial level 53m 25s100%2.8%/5hClaude Max 20x
Claude Opus 5claude-codeC69.0%level 3 · partial level 55m 48s100%0.8%/5hClaude Max 20x
Gemini 3.8 Flashantigravity-cliC56.1%level 2 · partial level 56m 2s100%0.9%/1wGoogle AI Pro
Grok 4.6grok-cliD53.5%level 3 · partial level 511m 26s100%0.2%/1wSuperGrok
GPT-5.6 Solcodex-cliD48.4%level 3 · partial level 53m 45s100%0%/1wChatGPT Pro $100
GPT-5.6 Terracodex-cliD45.8%level 2 · partial level 53m 1s100%0.2%/1wChatGPT Pro $100
Kimi K3 256Kkimi-cliD38.1%level 2 · partial level 57m 42s100%10.6%/5hKimi Allegretto
Composer 2.5cursor-cliD36.8%level 1 · partial level 52m 6s100%0.1%/1mCursor Pro
GPT-5.6 Lunacodex-cliF34.2%level 1 · partial level 53m 26s100%0%/1wChatGPT Pro $100
Claude Sonnet 5claude-codeF31.6%level 1 · partial level 54m 31s100%0.6%/5hClaude Max 20x
Gemini 3.1 Proantigravity-cliF23.2%none5m 16s100%1.1%/1wGoogle AI Pro
Claude Haiku 4.5claude-codeF21.3%none5m 0s100%0.4%/5hClaude Max 20x

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.