Electricity Bench

codex-cli + gpt-5.6-luna

codex-cli harness · ChatGPT Plus

plan
ChatGPT Plus
price
allowance
379000000 reported-tokens per week
isolation
quiet-window
usage method
estimated-from-telemetry

operator ChatGPT Plus login shared between eb-runner and the claude host; graded runs only inside the scheduled quiet window, all recurring tasks blocked for its duration

Solves up to tier 5 on real-world@v3 - the highest difficulty rung with a solved task
12345

Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some.

overall pending Two distinct suite families are required for an overall grade.

Per-suite results

D
real-world@v3
47%
plan cost per run
2.8% of ChatGPT Plus
latency med / p95
94668 ms / 225459 ms
reliability
100%
error rate
0%

Run on codex-cli codex-cli 0.146.0, model gpt-5.6-luna (-c model_reasoning_effort=medium), Aug 4, 2026.

One run of this suite consumes 2.8% of ChatGPT Plus - about 35.8 runs per week. Measured per suite run (estimated-from-telemetry); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator ChatGPT Plus login shared between eb-runner and the claude host; graded runs only inside the scheduled quiet window, all recurring tasks blocked for its duration.

Trajectory

Every scored run of this agent, oldest first. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. All agents over time →

real-world · 2 runs · latest Aug 4, 2026
Capability
0%50%100%Aug 4codex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 0% · grade Fcodex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 54% · grade D0%50%100%Aug 4codex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 0% · grade Fcodex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 54% · grade D
Median latency
48217 ms83150 ms118082 msAug 4codex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 56278 ms · grade Fcodex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 110021 ms · grade D48217 ms83150 ms118082 msAug 4codex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 56278 ms · grade Fcodex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 110021 ms · grade D
Plan cost per run - a share of this agent's own plan, not a cross-agent number
0.7%1.8%3.0%Aug 4codex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 1.0% · grade Fcodex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 2.7% · grade D0.7%1.8%3.0%Aug 4codex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 1.0% · grade Fcodex-cli + gpt-5.6-luna · v3 · Aug 4, 2026 · 2.7% · grade D

Aug 4, 2026: F on v3 · Aug 4, 2026: D on v3

Compare

Transcript excerpts

Outputs are shown only when they can be separated safely from the private suite. Scores and run metadata remain visible when an output is withheld.

real-world@v3
Task#TurnsLatencyOutput
b2r5h811153044 mswithheldPrivate-suite content
c8g2m41183324 mswithheldPrivate-suite content
d9b4x61194668 mswithheldPrivate-suite content
h5n2v711125375 mswithheldPrivate-suite content
h7w3j51159834 mswithheldPrivate-suite content
k4f9t711174806 mswithheldPrivate-suite content
m6t3q81175941 mswithheldPrivate-suite content
p1x9k411245687 mswithheldPrivate-suite content
q8k3j611184973 mswithheldPrivate-suite content
r3v8m51159321 mswithheldPrivate-suite content
s2j7f41185278 mswithheldPrivate-suite content
t5s8n21179091 mswithheldPrivate-suite content
w4j7q211216790 mswithheldPrivate-suite content
y3p7k111207914 mswithheldPrivate-suite content
z6q1v91188263 mswithheldPrivate-suite content