GPT-6 Luna scores D overall on Electricity Bench and clears tier 2 of 5 on real-world issues. Its best suite is spec planning at C-. One pass of our task set uses 1% of ChatGPT Plus's weekly limit. In our own use it is our pick for high-volume work: the smallest problems.Measured
Clears tier 2 of 5 (partial at tiers 3, 4 and 5) on Real-world issues
Capability mean across 3 graded suite families.
capability
50%
plan cost per run
1.0% of ChatGPT Plus
median task time
1m 0s
reliability
100%
12345
What the tiers mean
"Clears" marks the highest rung where every task at it and below was solved. Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some. The headline stops at the last unbroken rung, so a solved tier above a half-solved one counts only as partial.
Clears level 1 of 5 (partial at levels 2, 3, 4 and 5) on Vibe coding v1 - the vaguest report level at which it still fixed every bug, having also cleared every richer level below
12345
What the levels mean
Level 1: a full brief naming the surface and the acceptance · level 3: the report as filed · level 5: a vague vibe report in a non-technical user's words. Context falls as the level rises, so vaguer is harder. A cleared level fixed every bug at it and below; a partial level fixed some. The headline stops at the last unbroken level, so a fix above a half-cleared level counts only as partial.
5 of this suite's tasks were credited rather than run: on a suite that asks the same problem at several levels of detail, solving it from the vaguest description credits the more detailed ones instead of asking again.
10 of 15 tasks solved - grouped by tier, easiest first; hover a square for its time and turns. solved partial not solved
Run on codex-cli codex-cli 0.155.1, model gpt-6-luna (-c model_reasoning_effort=medium), week 2026-W39.
One run of this suite consumes 1.0% of ChatGPT Plus's 5h window and 1.0% of the weekly limit.
ChatGPT Plus's weekly limit holds about 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Published at 6.0 while the native sample is two session windows deep; the codex meter only began reporting both windows on 2026-08-31.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a shared personal subscription during a quiet window, not a dedicated bench login.
0 of 4 tasks solved - hover a square for its time and turns. solved partial not solved
Run on codex-cli codex-cli 0.155.1, model gpt-6-luna (-c model_reasoning_effort=medium), week 2026-W39.
One run of this suite consumes under 1% of ChatGPT Plus's weekly limit. The meter reports whole percent and did not move over the run, so this is the resolution's ceiling, not a measured share - at the plan's listed price, at most $0.011 per task.
ChatGPT Plus's weekly limit holds about 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Published at 6.0 while the native sample is two session windows deep; the codex meter only began reporting both windows on 2026-08-31.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a shared personal subscription during a quiet window, not a dedicated bench login.
13 of 25 tasks solved (5 credited without running) - each row is one bug, vaguest report first and context growing to the right; hover a square for its time and turns. solved partial not solved credited
Run on codex-cli codex-cli 0.155.1, model gpt-6-luna (-c model_reasoning_effort=medium), week 2026-W39.
ChatGPT Plus's weekly limit holds about 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Published at 6.0 while the native sample is two session windows deep; the codex meter only began reporting both windows on 2026-08-31.
Task outputs from these runs are withheld: publishing them would reveal the private suite. Per-task scores, times and turns are in the chips above.
Trajectory
One point per grading week - a re-run within a week shows only its latest execution. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. Click a point or week for plan usage. All agents over time →
Put GPT-6 Luna's current grade in a README or a docs page.
Markdown
[](https://electricitybench.com/models/codex-cli-gpt-6-luna/)
HTML
<a href="https://electricitybench.com/models/codex-cli-gpt-6-luna/"><img src="https://electricitybench.com/badge/codex-cli-gpt-6-luna.svg" alt="GPT-6 Luna on Electricity Bench: D"></a>
Badges update every Monday with the weekly run. GitHub renders them from this URL.
Frequently asked questions
Is GPT-6 Luna worth it on ChatGPT Plus?
On our private real-world suite (real-world@v3), GPT-6 Luna earned grade D+ with 50% capability, clearing every problem up to tier 2 of 5 - a bug spanning a few files and landing some tier 5 problems. It runs on ChatGPT Plus at $20/month (public price as of 2026-07-20). One full suite run consumed 1.0% of the plan's week window. Measured 2026-09-23.
Is GPT-6 Luna good for coding?
GPT-6 Luna clears tier 2 of our five-tier real-world ladder - every task at that tier and below - where tier 1 is a localized single-file bug and tier 5 is a problem that took a human engineer hours. It solves some but not all tasks up at tier 5. Its capability score on real-world@v3 is 50% (grade D+). Last benchmarked 2026-09-23.
How much of ChatGPT Plus does one GPT-6 Luna task use?
One full suite run consumed 1.0% of ChatGPT Plus's weekly limit, read off the provider's own meter before and after the run. A benchmark task is a whole multi-turn agent session, not a message, and the share belongs to this plan alone: it is not comparable with another agent's plan. Measured 2026-09-23.