codex-cli + gpt-6-astra
codex-cli harness · ChatGPT Pro $100 ($100/mo)
Clears tier 2 of 5 (partial at tiers 3 and 5) on Real-world issues
Capability mean across 3 graded suite families.
What the tiers mean
"Clears" marks the highest rung where every task at it and below was solved. Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some. The headline stops at the last unbroken rung, so a solved tier above a half-solved one counts only as partial.
What the levels mean
Level 1: a full brief naming the surface and the acceptance · level 3: the report as filed · level 5: a vague vibe report in a non-technical user's words. Context falls as the level rises, so vaguer is harder. A cleared level fixed every bug at it and below; a partial level fixed some. The headline stops at the last unbroken level, so a fix above a half-cleared level counts only as partial.
10 of this suite's tasks were credited rather than run: on a suite that asks the same problem at several levels of detail, solving it from the vaguest description credits the more detailed ones instead of asking again.
Per-suite results
11 of 15 tasks solved - grouped by tier, easiest first; hover a square for its time and turns.
Run on codex-cli codex-cli 0.153.4, model gpt-6-astra (-c model_reasoning_effort=medium), week 2026-W36.
One run of this suite consumes 3.0% of ChatGPT Pro $100.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: RunByAI OpenAI contestant ChatGPT Pro $100 login, moved onto eb-runner on 2026-09-05 and still present on the claude host RunByAI pane; graded runs only inside the scheduled quiet window, all recurring tasks blocked for its duration.
0 of 4 tasks solved - hover a square for its time and turns.
Run on codex-cli codex-cli 0.153.4, model gpt-6-astra (-c model_reasoning_effort=medium), week 2026-W36.
One run of this suite consumes 1.0% of ChatGPT Pro $100.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: RunByAI OpenAI contestant ChatGPT Pro $100 login, moved onto eb-runner on 2026-09-05 and still present on the claude host RunByAI pane; graded runs only inside the scheduled quiet window, all recurring tasks blocked for its duration.
20 of 25 tasks solved (10 credited without running) - each row is one bug, vaguest report first and context growing to the right; hover a square for its time and turns.
Run on codex-cli codex-cli 0.153.4, model gpt-6-astra (-c model_reasoning_effort=medium), week 2026-W36.
One good solve on this suite consumes 0.2% of ChatGPT Pro $100 on average.
How the plan cost was measured
Counted over each issue's first solve (5 solves); the walk's failed harder cells and confirmation runs are the benchmark's own search cost and are excluded (full walk: 4.0% over 15 cells). Measured per cell run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: RunByAI OpenAI contestant ChatGPT Pro $100 login, moved onto eb-runner on 2026-09-05 and still present on the claude host RunByAI pane; graded runs only inside the scheduled quiet window, all recurring tasks blocked for its duration.
Task outputs from these runs are withheld: publishing them would reveal the private suite. Per-task scores, times and turns are in the chips above.
Frequently asked questions
Is GPT-6 Astra worth it on ChatGPT Pro $100?
On our private real-world suite (real-world@v3), GPT-6 Astra earned grade C with 57% capability, clearing every problem up to tier 2 of 5 - a bug spanning a few files and landing some tier 5 problems. It runs on ChatGPT Pro $100 at $100/month (public price as of 2026-09-05). One full suite run consumed 3.0% of the plan's week window. Measured 2026-09-05.
How good is GPT-6 Astra at real-world coding?
GPT-6 Astra clears tier 2 of our five-tier real-world ladder - every task at that tier and below - where tier 1 is a localized single-file bug and tier 5 is a problem that took a human engineer hours. It solves some but not all tasks up at tier 5. Its capability score on real-world@v3 is 57% (grade C). Last benchmarked 2026-09-05.
Compare
- GPT-6 Astra vs Gemini 3.1 Pro
- GPT-6 Astra vs Gemini 3.6 Flash
- GPT-6 Astra vs Gemini 3.7 Flash
- GPT-6 Astra vs Gemini 3.8 Flash
- GPT-6 Astra vs Claude Fable 5.1
- GPT-6 Astra vs Claude Fable 5
- GPT-6 Astra vs Claude Haiku 4.5
- GPT-6 Astra vs Claude Opus 5
- GPT-6 Astra vs Claude Sonnet 5
- GPT-6 Astra vs GPT-5.6 Luna
- GPT-6 Astra vs GPT-5.6 Sol
- GPT-6 Astra vs GPT-5.6 Terra
- GPT-6 Astra vs Composer 2.5
- GPT-6 Astra vs Grok 4.5
- GPT-6 Astra vs Grok 4.6
- GPT-6 Astra vs Kimi K3 256K