Electricity Bench

Claude Sonnet 5 vs Kimi K3 256K

Claude Sonnet 5 is no longer on the leaderboard. The grades below are the last ones measured, so read them as history rather than as this week's board. See the current roster.

Kimi K3 256K leads overall, D+ to D. Claude Sonnet 5 is about 6.0x cheaper per task. By suite, Kimi K3 256K takes real-world issues (C- to D+) and spec planning (C- to D). Take Claude Sonnet 5 anyway when you want the side that is cheaper per task.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: Kimi K3 256K vs Claude Sonnet 5

Kimi K3 256K leads overall, D+ to D. By suite: Kimi K3 256K takes real-world issues, C- to D+ (60% vs 53% capability); Kimi K3 256K takes spec planning, C- to D (56% vs 44% capability); vibe coding is even at F (31% vs 34% capability). Claude Sonnet 5 used 2% of Claude Max 5x's weekly limit per pass, $0.024 per task at the listed price. Kimi K3 256K used 31% of Kimi Allegretto's weekly limit per pass, $0.146 per task at the listed price. Per task, Claude Sonnet 5 is about 6.0x cheaper ($0.024 vs $0.146). Kimi K3 256K is faster on real-world issues: a median task takes 3m 31s against 3m 35s. Both finished every task on it. Reasons to pick Claude Sonnet 5 anyway: it is cheaper per task ($0.024 vs $0.146).

Changed since W39: Claude Sonnet 5: vibe coding D+ to F; Kimi K3 256K: real-world issues usage 23% to 21% of the weekly limit, spec planning D+ to C-, spec planning usage 11% to 10% of the weekly limit, vibe coding usage 3% to 2.6% of the weekly limit.

Claude Sonnet 5Clears tier 3 (partial tier 5) at 12.0% of Claude Max 5x per runDOVERALLKimi K3 256KClears tier 4 (partial tier 5) at 104.0% of Kimi Allegretto per runD+OVERALL
real-world@v3D+vsC-
DimensionClaude Sonnet 5Kimi K3 256KAdvantage
Capability53%60%+7% (Kimi K3 256K leads)
Median task time3m 35s3m 31s4s faster (Kimi K3 256K leads)
Approx. task cost$0.015 / task$0.126 / task$0.110 / task lower (Claude Sonnet 5 leads)

Claude Sonnet 5: 12.0% of Claude Max 5x (12.0% / 5h, 1.0% / 1w) per run. Its weekly limit holds about 13.8 of the week's 33.6 5h windows. Kimi K3 256K: 104.0% of Kimi Allegretto (104.0% / 5h, 21.0% / 1w) per run. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

Work per window: Claude Sonnet 5 1218 vs Kimi K3 256K 140 suite runs per month (Claude Sonnet 5 leads); per plan-dollar: 12.18 vs 3.60 (Claude Sonnet 5 leads). Work figures compare outputs, so unlike quota shares they are comparable across plans.

spec-planning@v1DvsC-
DimensionClaude Sonnet 5Kimi K3 256KAdvantage
Capability44%56%+12% (Kimi K3 256K leads)
Median task time5m 12s5m 26s14s faster (Claude Sonnet 5 leads)
Approx. task cost$0.057 / task$0.224 / task$0.167 / task lower (Claude Sonnet 5 leads)

Claude Sonnet 5: 5.0% of Claude Max 5x (5.0% / 5h, 1.0% / 1w) per run. Its weekly limit holds about 13.8 of the week's 33.6 5h windows. Kimi K3 256K: 52.0% of Kimi Allegretto (52.0% / 5h, 10.0% / 1w) per run. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

Work per window: Claude Sonnet 5 2922 vs Kimi K3 256K 281 suite runs per month (Claude Sonnet 5 leads); per plan-dollar: 29.22 vs 7.20 (Claude Sonnet 5 leads). Work figures compare outputs, so unlike quota shares they are comparable across plans.

vibe-coding@v1FvsF
DimensionClaude Sonnet 5Kimi K3 256KAdvantage
Capability31%34%+3% (Kimi K3 256K leads)
Median task time4m 31s7m 33s3m 2s faster (Claude Sonnet 5 leads)
Approx. task cost$0.007 / task$0.233 / task$0.226 / task lower (Claude Sonnet 5 leads)

Claude Sonnet 5: 1.0% of Claude Max 5x (1.0% / 5h) per good solve. Its weekly limit holds about 13.8 of the week's 33.6 5h windows. Kimi K3 256K: 13.2% of Kimi Allegretto (13.2% / 5h, 2.6% / 1w) per good solve. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

Work per window: Claude Sonnet 5 14610 vs Kimi K3 256K 1107 good solves per month (Claude Sonnet 5 leads); per plan-dollar: 146.10 vs 28.38 (Claude Sonnet 5 leads). Work figures compare outputs, so unlike quota shares they are comparable across plans.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, Claude Sonnet 5 or Kimi K3 256K?

Kimi K3 256K clears higher: tier 4 of 5 versus tier 3 for Claude Sonnet 5 on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-09-28.

Which is cheaper per task, Claude Sonnet 5 or Kimi K3 256K?

Claude Sonnet 5 used 2% of Claude Max 5x's weekly limit per pass, $0.024 per task at the listed price. Kimi K3 256K used 31% of Kimi Allegretto's weekly limit per pass, $0.146 per task at the listed price. Per task, Claude Sonnet 5 is about 6.0x cheaper ($0.024 vs $0.146).

Which gives more coding work per dollar, Claude Sonnet 5 or Kimi K3 256K?

Claude Sonnet 5 delivers more benchmark work per subscription dollar: 12.18 suite runs per plan-dollar versus 3.60 for Kimi K3 256K (at $100/month priced 2026-07-20, and $39/month priced 2026-07-20). Work figures compare outputs, so unlike raw quota shares they are comparable across plans.

What do Claude Sonnet 5 and Kimi K3 256K cost a month?

Claude Sonnet 5 runs on Claude Max 5x at $100/month (public price as of 2026-07-20). Kimi K3 256K runs on Kimi Allegretto at $39/month (public price as of 2026-07-20).