Electricity Bench

GPT-5.6 Luna

codex-cli harness · ChatGPT Pro 5x ($100/mo)

Coding agent review, updated weekly (latest W40 2026).

GPT-5.6 Luna scores D overall on Electricity Bench and clears tier 1 of 5 on real-world issues. Its best suite is spec planning at C-. One pass of our task set uses <1% of ChatGPT Pro 5x's weekly limit. Every grade comes from private tasks the agent has not seen, run headless on the subscription you would actually buy.Measured

Superseded by GPT-6 Luna (week 2026-W39). These grades stay published as the record of what this agent scored while it ran.

DOVERALL

Clears tier 1 of 5 (partial at tiers 2 and 5) on Real-world issues

Capability mean across 3 graded suite families.

capability
45%
plan cost per run
<1% of ChatGPT Pro 5x's weekly limit
median task time
3m 7s
reliability
100%
12345
What the tiers mean

"Clears" marks the highest rung where every task at it and below was solved. Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some. The headline stops at the last unbroken rung, so a solved tier above a half-solved one counts only as partial.

Clears level 1 of 5 (partial at levels 2, 3, 4 and 5) on Vibe coding v1 - the vaguest report level at which it still fixed every bug, having also cleared every richer level below
12345
What the levels mean

Level 1: a full brief naming the surface and the acceptance · level 3: the report as filed · level 5: a vague vibe report in a non-technical user's words. Context falls as the level rises, so vaguer is harder. A cleared level fixed every bug at it and below; a partial level fixed some. The headline stops at the last unbroken level, so a fix above a half-cleared level counts only as partial.

3 of this suite's tasks were credited rather than run: on a suite that asks the same problem at several levels of detail, solving it from the vaguest description credits the more detailed ones instead of asking again.

Per-suite results

D
Real-world issues v3
45%
plan cost per run
<1% of ChatGPT Pro 5x's weekly limit
median task time
3m 7s
reliability
100%
tier 1tier 2tier 3tier 4tier 5

10 of 15 tasks solved - grouped by tier, easiest first; hover a square for its time and turns.

Run on codex-cli codex-cli 0.155.1, model gpt-5.6-luna (-c model_reasoning_effort=medium), week 2026-W39.

One run of this suite consumes under 1% of ChatGPT Pro 5x's weekly limit. The meter reports whole percent and did not move over the run, so this is the resolution's ceiling, not a measured share - at the plan's listed price, at most $0.015 per task.

How the plan cost was measured

Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a shared personal subscription during a quiet window, not a dedicated bench login.

How the limits work on Codex CLI: /limits/codex-cli/

C-
Spec planning v1
57%
plan cost per run
<1% of ChatGPT Pro 5x's weekly limit
median task time
3m 25s
reliability
100%

0 of 4 tasks solved - hover a square for its time and turns.

Run on codex-cli codex-cli 0.155.1, model gpt-5.6-luna (-c model_reasoning_effort=medium), week 2026-W39.

One run of this suite consumes under 1% of ChatGPT Pro 5x's weekly limit. The meter reports whole percent and did not move over the run, so this is the resolution's ceiling, not a measured share - at the plan's listed price, at most $0.057 per task.

How the plan cost was measured

Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a shared personal subscription during a quiet window, not a dedicated bench login.

How the limits work on Codex CLI: /limits/codex-cli/

F
Vibe coding v1
25%
plan cost per solve
<1% of ChatGPT Pro 5x's weekly limit
median task time
3m 40s
reliability
100%
level 5level 4level 3level 2level 1issue aissue bissue cissue dissue e

13 of 25 tasks solved (3 credited without running) - each row is one bug, vaguest report first and context growing to the right; hover a square for its time and turns.

Run on codex-cli codex-cli 0.155.1, model gpt-5.6-luna (-c model_reasoning_effort=medium), week 2026-W39.

One good solve on this suite consumes under 1% of ChatGPT Pro 5x's weekly limit. The meter reports whole percent and did not move over the run, so this is the resolution's ceiling, not a measured share - at the plan's listed price, at most $0.230 per task.

How the plan cost was measured

Counted over each issue's first solve (5 solves); the walk's failed harder cells and confirmation runs are the benchmark's own search cost and are excluded. Measured per cell run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a shared personal subscription during a quiet window, not a dedicated bench login.

How the limits work on Codex CLI: /limits/codex-cli/

Task outputs from these runs are withheld: publishing them would reveal the private suite. Per-task scores, times and turns are in the chips above.

Trajectory

One point per grading week - a re-run within a week shows only its latest execution. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. Click a point or week for plan usage. All agents over time →

Suite
Real-world issues · 8 weeks graded · latest 2026-W39
30%40%50%60%2026-W322026-W342026-W362026-W382026-W39GPT-5.6 Luna · 2026-W32 · Aug 4, 2026 · 47% · DGPT-5.6 Luna · 2026-W33 · Aug 10, 2026 · 51% · D+GPT-5.6 Luna · 2026-W34 · Aug 17, 2026 · 53% · D+GPT-5.6 Luna · 2026-W35 · Aug 23, 2026 · 47% · DGPT-5.6 Luna · 2026-W36 · Sep 1, 2026 · 42% · DGPT-5.6 Luna · 2026-W37 · Sep 6, 2026 · 53% · D+GPT-5.6 Luna · 2026-W38 · Sep 14, 2026 · 47% · DGPT-5.6 Luna · 2026-W39 · Sep 21, 2026 · 45% · D

Weekly heads

Newest week first. Click a row for usage.

WeekGradeCapabilityClearsTask timeUsage /1w
2026-W39 · v3nowD45%tier 1 · p53m 7s0.0% of ChatGPT Pro 5x
2026-W38 · v3D47%tier 4 · p51m 50s0.0% of ChatGPT Pro 5x
2026-W37 · v3D+53%tier 0 · p52m 12s0.0% of ChatGPT Pro 5x
2026-W36 · v3D42%tier 1 · p51m 53s0.2% of ChatGPT Pro 5x
2026-W35 · v3D47%tier 4 · p52m 5s0.2% of ChatGPT Pro 5x
2026-W34 · v3D+53%tier 3 · p51m 59s0.2% of ChatGPT Pro 5x
2026-W33 · v3D+51%tier 0 · p52m 15s0.2% of ChatGPT Pro 5x
2026-W32 · v3D47%tier 4 · p51m 35s0.1% of ChatGPT Pro 5x

Embed this grade

Put GPT-5.6 Luna's current grade in a README or a docs page.

GPT-5.6 Luna on Electricity Bench: D

Markdown
[![GPT-5.6 Luna on Electricity Bench: D](https://electricitybench.com/badge/codex-cli-gpt-5-6-luna.svg)](https://electricitybench.com/models/codex-cli-gpt-5-6-luna/)
HTML
<a href="https://electricitybench.com/models/codex-cli-gpt-5-6-luna/"><img src="https://electricitybench.com/badge/codex-cli-gpt-5-6-luna.svg" alt="GPT-5.6 Luna on Electricity Bench: D"></a>

Badges update every Monday with the weekly run. GitHub renders them from this URL.

Frequently asked questions

Is GPT-5.6 Luna worth it on ChatGPT Pro 5x?

On our private real-world suite (real-world@v3), GPT-5.6 Luna earned grade D with 45% capability, clearing every problem up to tier 1 of 5 - a localized single-file bug and landing some tier 5 problems. It runs on ChatGPT Pro 5x at $100/month (public price as of 2026-09-05). One full suite run consumed 0.0% of the plan's week window. Measured 2026-09-21.

Is GPT-5.6 Luna good for coding?

GPT-5.6 Luna clears tier 1 of our five-tier real-world ladder - every task at that tier and below - where tier 1 is a localized single-file bug and tier 5 is a problem that took a human engineer hours. It solves some but not all tasks up at tier 5. Its capability score on real-world@v3 is 45% (grade D). Last benchmarked 2026-09-21.

How much of ChatGPT Pro 5x does one GPT-5.6 Luna task use?

One full suite run consumed 0.0% of ChatGPT Pro 5x's weekly limit, read off the provider's own meter before and after the run. A benchmark task is a whole multi-turn agent session, not a message, and the share belongs to this plan alone: it is not comparable with another agent's plan. Measured 2026-09-21.

Compare