Electricity Bench

2026-W38 grading run

One public snapshot for the week, assembled as each suite publishes. Grades compare only within the same suite version.

Week 38 of 2026 graded 13 coding agents across 3 published suites. Claude Opus 5 leads real-world issues with C+. The biggest move since W37 2026 is GPT-5.6 Luna on vibe coding, F to C-. Grades compare only within the same suite version, and plan usage is each agent's share of its own subscription.Measured

3
published suites
13
graded agents
Sep 14, 2026
latest result

Suite snapshots

Real-world issues v3

13 graded agents

Lead: Claude Opus 5 · C+ · 73.3%

Published through Sep 14, 2026
Spec planning v1

13 graded agents

Lead: GPT-6 Astra · C+ · 74.4%

Published through Sep 14, 2026
Vibe coding v1

13 graded agents

Lead: GPT-6 Astra · C+ · 74.2%

Published through Sep 14, 2026

A suite can appear here before the others. Rebuilding after another suite publishes adds its snapshot without changing this URL.