Electricity Bench

GPT-6 Sol is a step back from 5.6, Luna stands still

Model Release

By the Electricity Bench editorial team · Drafted with AI, reviewed and approved by a human editor

Graded two days after release: GPT-6 Sol fixes less than GPT-5.6 Sol on well under half the usage, and GPT-6 Luna is faster but not better.

OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, both at half the API price or less of the GPT-5.6 models they replace: $2 and $10 per million input and output tokens for Sol, $0.10 and $0.50 for Luna. OpenAI says Sol at max effort nearly matches Claude Fable 5 on its software engineering benchmark at a fifth of the cost, and that Luna at max effort lands near Claude Opus 5. We graded both as the successors to GPT-5.6 Sol and GPT-5.6 Luna.

What the results show

Our take: neither model is worth switching to for coding, and that matches our own use. GPT-5.6 Sol is still in Codex and stays the better OpenAI pick. With this release a Claude subscription becomes the clear choice even at $20, without needing Fable, because Opus 5.5 fixes far more for about the same weekly usage. GPT-6 works differently from the GPT-5.6 models under the hood and Codex takes time to adjust to a new model line, so we expect both to improve over the coming weeks. We are still waiting on results for Grok 4.7, but from the chatter online it does not seem to perform that well.

Grades are comparable within the same suite version. See the full GPT-6 Sol and GPT-6 Luna scorecards or follow the field over time on Trends.