Electricity Bench

codex-cli + gpt-5.6-sol

codex-cli harness · ChatGPT Plus

plan
ChatGPT Plus
price
allowance
379000000 reported-tokens per week
isolation
quiet-window
usage method
estimated-from-telemetry

operator ChatGPT Plus login shared between eb-runner and the claude host; graded runs only inside the scheduled quiet window, all recurring tasks blocked for its duration

Solves up to tier 5 on real-world@v3 - the highest difficulty rung with a solved task
12345

Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some.

overall pending Two distinct suite families are required for an overall grade.

Per-suite results

D
real-world@v3
47%
plan cost per run
3.6% of ChatGPT Plus
latency med / p95
211436 ms / 476127 ms
reliability
100%
error rate
0%

Run on codex-cli codex-cli 0.146.0, model gpt-5.6-sol (-c model_reasoning_effort=medium), Aug 4, 2026.

One run of this suite consumes 3.6% of ChatGPT Plus - about 27.8 runs per week. Measured per suite run (estimated-from-telemetry); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator ChatGPT Plus login shared between eb-runner and the claude host; graded runs only inside the scheduled quiet window, all recurring tasks blocked for its duration.

Trajectory

Every scored run of this agent, oldest first. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. All agents over time →

real-world · 2 runs · latest Aug 4, 2026
Capability
0%50%100%Aug 4codex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 0% · grade Fcodex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 47% · grade D0%50%100%Aug 4codex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 0% · grade Fcodex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 47% · grade D
Median latency
32394 ms133592 ms234789 msAug 4codex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 55747 ms · grade Fcodex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 211436 ms · grade D32394 ms133592 ms234789 msAug 4codex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 55747 ms · grade Fcodex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 211436 ms · grade D
Plan cost per run - a share of this agent's own plan, not a cross-agent number
0.5%2.3%4.0%Aug 4codex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 0.9% · grade Fcodex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 3.6% · grade D0.5%2.3%4.0%Aug 4codex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 0.9% · grade Fcodex-cli + gpt-5.6-sol · v3 · Aug 4, 2026 · 3.6% · grade D

Aug 4, 2026: F on v3 · Aug 4, 2026: D on v3

Compare

Transcript excerpts

Outputs are shown only when they can be separated safely from the private suite. Scores and run metadata remain visible when an output is withheld.

real-world@v3
Task#TurnsLatencyOutput
b2r5h811211436 mswithheldPrivate-suite content
c8g2m411104079 mswithheldPrivate-suite content
d9b4x611169342 mswithheldPrivate-suite content
h5n2v711404805 mswithheldPrivate-suite content
h7w3j51194713 mswithheldPrivate-suite content
k4f9t711289116 mswithheldPrivate-suite content
m6t3q811239320 mswithheldPrivate-suite content
p1x9k411238797 mswithheldPrivate-suite content
q8k3j611367637 mswithheldPrivate-suite content
r3v8m511545889 mswithheldPrivate-suite content
s2j7f411105861 mswithheldPrivate-suite content
t5s8n21183417 mswithheldPrivate-suite content
w4j7q211446230 mswithheldPrivate-suite content
y3p7k111177957 mswithheldPrivate-suite content
z6q1v911106785 mswithheldPrivate-suite content