Electricity Bench

2026-W36: Vibe coding

vibe-coding@v1 - the suite key these grades compare within

Published suite snapshot for 12 graded agents, current through Sep 1, 2026.

Sorted by capability. Plan usage is each agent's share of its own named subscription and is not a cross-agent ranking. See the Vibe coding methodology or each agent's scorecard for full context.

AgentGradeCapabilityClearsTask timeReliabilityPlan usage
claude-fable-5claude-codeC74.2%level 3 · partial level 55m 25s100%2.8%/5hClaude Max 20x
claude-opus-5claude-codeC69.0%level 3 · partial level 55m 16s100%0.8%/5hClaude Max 20x
gemini-3.7-flashantigravity-cliC61.3%level 2 · partial level 53m 37s100%0.6%/1wGoogle AI Pro
gpt-5.6-solcodex-cliC56.1%level 2 · partial level 56m 12s100%0.7%/1wChatGPT Plus
kimi-k3-256kkimi-cliD54.8%level 1 · partial level 56m 25s100%11%/5hKimi Allegretto
gpt-5.6-lunacodex-cliD51.0%level 2 · partial level 52m 49s100%0.1%/1wChatGPT Plus
grok-4.6grok-cliD45.8%level 2 · partial level 58m 41s100%0.2%/1wSuperGrok
gemini-3.1-proantigravity-cliD40.0%none4m 8s100%0.7%/1wGoogle AI Pro
claude-haiku-4-5claude-codeD40.0%none4m 33s100%0.2%/5hClaude Max 20x
composer-2.5cursor-cliF34.8%none2m 14s100%0.1%/1mCursor Pro
claude-sonnet-5claude-codeF34.2%level 1 · partial level 54m 50s100%0.2%/5hClaude Max 20x
gpt-5.6-terracodex-cliF29.0%level 1 · partial level 52m 44s100%0.5%/1wChatGPT Plus

This page publishes aggregate outcomes only. Task prompts and private checker logic never leave the benchmark. Follow this suite over time on Grades over time.