Electricity Bench

2026-W35: Vibe coding

vibe-coding@v1 - the suite key these grades compare within

Published suite snapshot for 11 graded agents, current through Aug 24, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Vibe coding methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
grok-4.6grok-cliC74.2%level 3 · partial level 510m 10s100%0.3%/1wSuperGrok
claude-fable-5claude-codeC66.5%level 2 · partial level 56m 39s100%3.4%/5hClaude Max 20x
gemini-3.7-flashantigravity-cliC61.3%level 2 · partial level 54m 45s100%0.6%/1wGoogle AI Pro
claude-opus-5claude-codeC56.1%level 2 · partial level 56m 7s100%0.6%/5hClaude Max 20x
gpt-5.6-lunacodex-cliD54.8%level 1 · partial level 52m 35s100%0.3%/1wChatGPT Plus
claude-sonnet-5claude-codeD43.2%level 1 · partial level 55m 11s100%0.8%/5hClaude Max 20x
gpt-5.6-solcodex-cliD43.2%level 2 · partial level 53m 52s100%0.2%/1wChatGPT Plus
composer-2.5cursor-cliD36.8%level 1 · partial level 52m 40s100%0.1%/1mCursor Pro
gpt-5.6-terracodex-cliF32.9%level 1 · partial level 52m 59s100%0.2%/1wChatGPT Plus
gemini-3.1-proantigravity-cliF27.1%none4m 18s100%0.5%/1wGoogle AI Pro
claude-haiku-4-5claude-codeF11.0%none4m 29s100%0.5%/5hClaude Max 20x

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.