Electricity Bench

Grok 4.5 vs Kimi K3 256K

Grok 4.5 was superseded by Grok 4.7. The grades below are the last ones measured, so read them as history rather than as this week's board. See the current roster.

Grok 4.5 leads overall, C- to D+. Grok 4.5 is about 4.6x cheaper per task. By suite, real-world issues is even at C- and Grok 4.5 takes vibe coding (C- to F). The data gives no reason to pick Kimi K3 256K over Grok 4.5 this week.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: Kimi K3 256K vs Grok 4.5

Grok 4.5 leads overall, C- to D+. By suite: real-world issues is even at C- (60% vs 60% capability); Grok 4.5 takes vibe coding, C- to F (59% vs 34% capability). Only Kimi K3 256K has run spec planning. Grok 4.5 used 6.9% of SuperGrok's weekly limit per pass, $0.032 per task at the listed price. Kimi K3 256K used 31% of Kimi Allegretto's weekly limit per pass, $0.146 per task at the listed price. Per task, Grok 4.5 is about 4.6x cheaper ($0.032 vs $0.146). Grok 4.5 is faster on real-world issues: a median task takes 3m 9s against 3m 31s. Both finished every task on it. The data gives no reason to pick Kimi K3 256K over Grok 4.5 this week.

Changed since the previous graded week: Grok 4.5: real-world issues usage 8.0% to 6.9% of the weekly limit; Kimi K3 256K: real-world issues usage 23% to 21% of the weekly limit, spec planning D+ to C-, spec planning usage 11% to 10% of the weekly limit, vibe coding usage 3% to 2.6% of the weekly limit.

Grok 4.5Clears tier 4 (partial tier 5) at 6.9% of SuperGrok per runC-OVERALLKimi K3 256KClears tier 4 (partial tier 5) at 104.0% of Kimi Allegretto per runD+OVERALL
real-world@v3C-vsC-
DimensionGrok 4.5Kimi K3 256KAdvantage
Capability60%60%Even
Median task time3m 9s3m 31s22s faster (Grok 4.5 leads)
Approx. task cost$0.032 / task$0.126 / task$0.094 / task lower (Grok 4.5 leads)

Grok 4.5: 6.9% of SuperGrok (6.9% / 1w) per run. Kimi K3 256K: 104.0% of Kimi Allegretto (104.0% / 5h, 21.0% / 1w) per run. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

spec-planning@v1-vsC-

Only Kimi K3 256K has run this suite, so there is no head-to-head yet.

vibe-coding@v1C-vsF
DimensionGrok 4.5Kimi K3 256KAdvantage
Capability59%34%+25% (Grok 4.5 leads)
Median task time3m 40s7m 33s3m 53s faster (Grok 4.5 leads)
Approx. task cost$0.031 / task$0.233 / task$0.203 / task lower (Grok 4.5 leads)

Grok 4.5: 0.4% of SuperGrok (0.4% / 1w) per good solve. Kimi K3 256K: 13.2% of Kimi Allegretto (13.2% / 5h, 2.6% / 1w) per good solve. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, Grok 4.5 or Kimi K3 256K?

They are even on our ladder: both clear tier 4 of 5 at 60% capability on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-08-10.

Which is cheaper per task, Grok 4.5 or Kimi K3 256K?

Grok 4.5 used 6.9% of SuperGrok's weekly limit per pass, $0.032 per task at the listed price. Kimi K3 256K used 31% of Kimi Allegretto's weekly limit per pass, $0.146 per task at the listed price. Per task, Grok 4.5 is about 4.6x cheaper ($0.032 vs $0.146).

What do Grok 4.5 and Kimi K3 256K cost a month?

Grok 4.5 runs on SuperGrok at $30/month (public price as of 2026-07-20). Kimi K3 256K runs on Kimi Allegretto at $39/month (public price as of 2026-07-20).