Electricity Bench

Opinion, not a scorecard

Our take: which agent we would actually use

The leaderboard reports what our private suites measure. This page is the other half of an honest answer to "which one should I buy?": the judgment the suites do not capture - how a model feels at speed, how good its harness is to live in, and how fast its plan really burns in day-to-day work.

This is our bias, stated openly

Everything below is opinion. It comes from our own experience using each model through its harness on its subscription - not from the graded runs. It is deliberately kept separate from the measured results: where a take disagrees with a published scorecard, the scorecard wins and we say so.

Last reviewed . Updated whenever our view changes, not on the weekly grading cadence.

The short version

  1. Claude Fable 5 - the strongest model available today, and the subscription makes it exceptional value.
  2. Claude Opus 5 - the daily driver: the allowance is effectively bottomless.
  3. Gemini 3.7 Flash - the likely heir to the fast lane, held back by its harness.
  4. Nothing else currently earns a place.

What we would use for what

Heavy lifting

the hardest problems

Claude Fable 5 (claude-code) is the strongest model you can run today. The subscription effectively subsidizes it, which makes the price-to-value ratio unmatched; whenever a problem is genuinely hard, this is our first choice.

Kimi K3 (kimi-cli) is the closest challenger on raw capability. It drains its allowance so quickly, though, that the plan is hard to recommend: the capability is real, the value is not.

Daily driver

unlimited everyday work

Claude Opus 5 (claude-code) is what we run all day. The allowance is generous enough that we have effectively never run out, which makes Opus the default for the steady stream of ordinary work.

Fast and cheap

UI work and minor tasks

This lane belonged to Grok 4.5 (grok-cli; retired): near-instant answers on a subscription that was nearly impossible to exhaust. Gemini 3.7 Flash (antigravity-cli) is the likely heir - 3.6 Flash was already remarkably fast for small tasks, and 3.7 is reported to be better - but the Antigravity harness trails the rest of the field badly enough to drag the whole product down.

No clear role right now

  • Grok 4.6 (grok-cli) - caught in the middle: no longer fast, and not strong enough to compete at the top.
  • GPT 5.6 Sol (codex-cli) - matched Fable at release, but has fallen well behind since. The generous limit resets read as a promotion; when they end, as we expect they will, the real cost of running Sol will show.
  • GPT 5.6 Luna (codex-cli) - has not been worth using since the pricing change.
  • GPT 5.6 Terra (codex-cli) - never made the case for choosing it over Opus.
  • Claude Sonnet 5 (claude-code) - slow and underwhelming: priced too close to Opus for the capability gap, and too heavy for the fast lane.
  • Claude Haiku 4.5 (claude-code) - a generation behind, and it shows.
  • Gemini 3.1 Pro (antigravity-cli) - overtaken by Opus 5, which leaves it without a job.
  • Composer 2.5 (cursor-cli) - an era behind the grok and flash class; not competitive.

For the measured side, start with the leaderboard, follow each agent over time on Trends, or read the methodology.