GPT-6 Sol scores D+ overall on Electricity Bench and clears tier 1 of 5 on real-world issues. Its best suite is spec planning at C. One pass of our task set uses 5% of ChatGPT Plus's weekly limit. In our own use it has no clear role right now.Measured
Clears tier 1 of 5 (partial at tiers 2, 4 and 5) on Real-world issues
Capability mean across 3 graded suite families.
capability
38%
plan cost per run
4.0% of ChatGPT Plus
median task time
2m 44s
reliability
100%
12345
What the tiers mean
"Clears" marks the highest rung where every task at it and below was solved. Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some. The headline stops at the last unbroken rung, so a solved tier above a half-solved one counts only as partial.
Clears level 2 of 5 (partial at levels 3, 4 and 5) on Vibe coding v1 - the vaguest report level at which it still fixed every bug, having also cleared every richer level below
12345
What the levels mean
Level 1: a full brief naming the surface and the acceptance · level 3: the report as filed · level 5: a vague vibe report in a non-technical user's words. Context falls as the level rises, so vaguer is harder. A cleared level fixed every bug at it and below; a partial level fixed some. The headline stops at the last unbroken level, so a fix above a half-cleared level counts only as partial.
9 of this suite's tasks were credited rather than run: on a suite that asks the same problem at several levels of detail, solving it from the vaguest description credits the more detailed ones instead of asking again.
9 of 15 tasks solved - grouped by tier, easiest first; hover a square for its time and turns. solved partial not solved
Run on codex-cli codex-cli 0.155.1, model gpt-6-sol (-c model_reasoning_effort=medium), week 2026-W39.
One run of this suite consumes 37.0% of ChatGPT Plus's 5h window and 4.0% of the weekly limit.
ChatGPT Plus's weekly limit holds about 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Published at 6.0 while the native sample is two session windows deep; the codex meter only began reporting both windows on 2026-08-31.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a shared personal subscription during a quiet window, not a dedicated bench login.
0 of 4 tasks solved - hover a square for its time and turns. solved partial not solved
Run on codex-cli codex-cli 0.155.1, model gpt-6-sol (-c model_reasoning_effort=medium), week 2026-W39.
One run of this suite consumes 8.0% of ChatGPT Plus's 5h window and 1.0% of the weekly limit.
ChatGPT Plus's weekly limit holds about 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Published at 6.0 while the native sample is two session windows deep; the codex meter only began reporting both windows on 2026-08-31.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a shared personal subscription during a quiet window, not a dedicated bench login.
19 of 25 tasks solved (9 credited without running) - each row is one bug, vaguest report first and context growing to the right; hover a square for its time and turns. solved partial not solved credited
Run on codex-cli codex-cli 0.155.1, model gpt-6-sol (-c model_reasoning_effort=medium), week 2026-W39.
One good solve on this suite consumes, on average, 2.4% of ChatGPT Plus's 5h window and 0.4% of the weekly limit.
ChatGPT Plus's weekly limit holds about 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Published at 6.0 while the native sample is two session windows deep; the codex meter only began reporting both windows on 2026-08-31.
How the plan cost was measured
Counted over each issue's first solve (5 solves); the walk's failed harder cells and confirmation runs are the benchmark's own search cost and are excluded (full walk: 8.0% over 16 cells). Measured per cell run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a shared personal subscription during a quiet window, not a dedicated bench login.
Task outputs from these runs are withheld: publishing them would reveal the private suite. Per-task scores, times and turns are in the chips above.
Trajectory
One point per grading week - a re-run within a week shows only its latest execution. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. Click a point or week for plan usage. All agents over time →
Put GPT-6 Sol's current grade in a README or a docs page.
Markdown
[](https://electricitybench.com/models/codex-cli-gpt-6-sol/)
HTML
<a href="https://electricitybench.com/models/codex-cli-gpt-6-sol/"><img src="https://electricitybench.com/badge/codex-cli-gpt-6-sol.svg" alt="GPT-6 Sol on Electricity Bench: D+"></a>
Badges update every Monday with the weekly run. GitHub renders them from this URL.
Frequently asked questions
Is GPT-6 Sol worth it on ChatGPT Plus?
On our private real-world suite (real-world@v3), GPT-6 Sol earned grade D- with 38% capability, clearing every problem up to tier 1 of 5 - a localized single-file bug and landing some tier 5 problems. It runs on ChatGPT Plus at $20/month (public price as of 2026-07-20). One full suite run consumed 4.0% of the plan's week window. Measured 2026-09-23.
Is GPT-6 Sol good for coding?
GPT-6 Sol clears tier 1 of our five-tier real-world ladder - every task at that tier and below - where tier 1 is a localized single-file bug and tier 5 is a problem that took a human engineer hours. It solves some but not all tasks up at tier 5. Its capability score on real-world@v3 is 38% (grade D-). Last benchmarked 2026-09-23.
How much of ChatGPT Plus does one GPT-6 Sol task use?
One full suite run consumed 4.0% of ChatGPT Plus's weekly limit, read off the provider's own meter before and after the run. A benchmark task is a whole multi-turn agent session, not a message, and the share belongs to this plan alone: it is not comparable with another agent's plan. Measured 2026-09-23.