Electricity Bench

Grok 4.5 vs Grok 4.6

Grok 4.5 was superseded by Grok 4.7. Grok 4.6 was superseded by Grok 4.7. The grades below are the last ones measured, so read them as history rather than as this week's board. See the current roster.

Both grade C- overall. Grok 4.5 is about 5.0x cheaper per task. By suite, Grok 4.5 takes real-world issues (C- to D) and vibe coding is even at C-. Both agents run the same private suite cut headless through their own tools, each on its own subscription.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: Grok 4.6 vs Grok 4.5

Both grade C- overall. By suite: Grok 4.5 takes real-world issues, C- to D (60% vs 47% capability); vibe coding is even at C- (59% vs 61% capability). Only Grok 4.6 has run spec planning. Grok 4.5 used 6.9% of SuperGrok's weekly limit per pass, $0.032 per task at the listed price. Grok 4.6 used 13% of SuperGrok Plus's weekly limit per pass, $0.157 per task at the listed price. Per task, Grok 4.5 is about 5.0x cheaper ($0.032 vs $0.157). Grok 4.5 is faster on real-world issues: a median task takes 3m 9s against 7m 10s. Both finished every task on it.

Changed since the previous graded week: Grok 4.5: real-world issues usage 8.0% to 6.9% of the weekly limit; Grok 4.6: real-world issues C- to D, real-world issues usage 4% to 3% of the weekly limit, vibe coding C+ to C-.

Grok 4.5Clears tier 4 (partial tier 5) at 6.9% of SuperGrok per runC-OVERALLGrok 4.6Clears tier 4 (partial tier 5) at 3.0% of SuperGrok Plus per runC-OVERALL
real-world@v3C-vsD
DimensionGrok 4.5Grok 4.6Advantage
Capability60%47%+13% (Grok 4.5 leads)
Median task time3m 9s7m 10s4m 1s faster (Grok 4.5 leads)
Approx. task cost$0.032 / task$0.046 / task$0.014 / task lower (Grok 4.5 leads)

Grok 4.5: 6.9% of SuperGrok (6.9% / 1w) per run. Grok 4.6: 3.0% of SuperGrok Plus (3.0% / 1w) per run. Each figure is a share of that agent's own plan, so there is no delta to draw.

spec-planning@v1-vsC

Only Grok 4.6 has run this suite, so there is no head-to-head yet.

vibe-coding@v1C-vsC-
DimensionGrok 4.5Grok 4.6Advantage
Capability59%61%+3% (Grok 4.6 leads)
Median task time3m 40s11m 14s7m 34s faster (Grok 4.5 leads)
Approx. task cost$0.031 / task<$0.230 / task$0.199 / task lower (Grok 4.5 leads)

Grok 4.5: 0.4% of SuperGrok (0.4% / 1w) per good solve. Grok 4.6: under 1% of SuperGrok Plus per good solve. Each figure is a share of that agent's own plan, so there is no delta to draw.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, Grok 4.5 or Grok 4.6?

Both clear tier 4 of 5, but Grok 4.5 scores higher on capability: 60% versus 47% on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-08-10.

Which is cheaper per task, Grok 4.5 or Grok 4.6?

Grok 4.5 used 6.9% of SuperGrok's weekly limit per pass, $0.032 per task at the listed price. Grok 4.6 used 13% of SuperGrok Plus's weekly limit per pass, $0.157 per task at the listed price. Per task, Grok 4.5 is about 5.0x cheaper ($0.032 vs $0.157).

What do Grok 4.5 and Grok 4.6 cost a month?

Grok 4.5 runs on SuperGrok at $30/month (public price as of 2026-07-20). Grok 4.6 runs on SuperGrok Plus at $100/month (public price as of 2026-09-14).