Electricity Bench

Opinion, not a scorecard

Our Take On Models

The leaderboard reports what our private suites measure. This page is the other half of an honest answer to "which one should I buy?": the judgment the suites do not capture - how a model feels at speed, how good its harness is to live in, and how fast its plan really burns in day-to-day work.

This is our bias, stated openly

Everything below is opinion. It comes from our own experience using each model through its harness on its subscription - not from the graded runs. It is deliberately kept separate from the measured results: where a take disagrees with a published scorecard, the scorecard wins and we say so.

Last reviewed . Updated whenever our view changes, not on the weekly grading cadence.

The short version

  1. Claude Opus 5.5 - the model we reach for first, for daily work and most hard problems. It is light on the plan.
  2. Claude Fable 5.1 - still the strongest on the hardest problems. With Opus 5.5 this good, we need it less often.
  3. Claude Sonnet 5.5 - the model for smaller things and simple tasks. Fast, and very light on the plan.

What we would use for what

Heavy lifting

the hardest problems

Claude Fable 5.1 (claude-code) is the strongest model you can run today. It uses far more of the plan than Opus 5.5, so we save it for the problems where Opus 5.5 gets stuck.

Claude Opus 5.5 (claude-code) takes the second slot. In our use it does most of what Fable does on a fraction of the plan usage, so it is what we reach for on most hard problems.

Daily work

the default for most tasks

Claude Opus 5.5 (claude-code) is our daily driver for a steady stream of small and medium tasks. Its allowance feels effectively bottomless, and Claude Code makes it easy to keep several pieces of work moving without losing track of what each agent is doing.

High-volume work

the smallest problems

Claude Sonnet 5.5 (claude-code) is now our model for smaller things and simple tasks. Give it a clearly defined task and it does the work fast on very little of the plan. We already use it on our own project and are extremely happy with it. Leave the problem open and Opus 5.5 is still the one to use.

GPT-6 Luna (codex-cli) is the alternative when cost matters most. Its native ChatGPT plan rate uses about one-fourteenth as much quota as Sol for the same token count.

No clear role right now

Models listed roughly by how capable they feel in our use, strongest first.

  • GPT-6 Sol (codex-cli) - GPT 5.6 Sol was our second pick for hard work because it felt faster than Fable. Opus 5.5 took that slot. It does more of the work on less of the plan. GPT-6 Sol looks like a step back.
  • GPT-6 Astra (codex-cli) - might be a step up on paper, but in our use and on the bench it does the same work as Sol, only faster and with a heavier plan burn, at two and a half times Sol's API price. Sol is the better value for everyday work and Fable the pick for hard problems, so Astra has no slot yet. Codex is not tuned around it, so this may change as the harness catches up.
  • Grok 4.7 (grok-cli) - Grok 4.5 had potential because it was fast and capable enough for everyday work. Grok 4.6 went in the wrong direction and 4.7 goes further down it. It seems like xAI just switched tiers: we had to downgrade from high to medium effort, and at medium the results are the same as 4.6, while high takes way too long for any work. Too slow for routine work and not capable enough for difficult work, so no reason to use it.
  • GPT 5.6 Terra (codex-cli) - sits awkwardly between Sol and Luna. Sol is better for difficult work, while Luna is far cheaper for routine tasks. Terra does not offer enough over either one to justify choosing it, and it also loses to Opus as a general-purpose daily driver.
  • Kimi K3 (kimi-cli) - competed with Fable and Sol when it launched, but has regressed since. It now does noticeably less with the same prompts and still drains its allowance quickly.
  • Gemini 3.8 Flash (antigravity-cli) - fast and cheap on paper, but it takes about twice as long as its predecessor and burns far more of the plan doing it. Sonnet 5.5 does the small, simple work better. We do not recommend it. That may change as Antigravity improves around it.
  • Gemini 3.1 Pro (antigravity-cli) - too slow to compete with Flash for everyday work and too weak to challenge Opus or Fable on difficult problems. Neither Gemini model has a slot for us right now.
  • Composer 2.5 (cursor-cli) - not capable enough for serious work. It struggles with tasks that stronger models handle easily, and being cheap is its only advantage. Cursor may still be worth choosing for its IDE and access to several providers, but Composer itself is not a reason to buy it.
  • Claude Haiku 4.5 (claude-code) - terrible by current standards. It is too weak to trust with real work and has no speed or cost advantage that makes up for it. Sonnet 5.5 and Luna are dramatically better choices for fast, cheap tasks.

For the measured side, start with the leaderboard, follow each agent over time on Trends, or read the methodology.