Electricity Bench

Gemini 3.1 Pro vs Grok 4.5

Grok 4.5 was superseded by Grok 4.7. The grades below are the last ones measured, so read them as history rather than as this week's board. See the current roster.

Grok 4.5 leads overall, C- to F. Grok 4.5 is about 1.2x cheaper per task. By suite, Grok 4.5 takes real-world issues (C- to D+) and vibe coding (C- to F). The data gives no reason to pick Gemini 3.1 Pro over Grok 4.5 this week.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: Grok 4.5 vs Gemini 3.1 Pro

Grok 4.5 leads overall, C- to F. By suite: Grok 4.5 takes real-world issues, C- to D+ (60% vs 53% capability); Grok 4.5 takes vibe coding, C- to F (59% vs 27% capability). Only Gemini 3.1 Pro has run spec planning. Gemini 3.1 Pro used 15.9% of Google AI Pro's weekly limit per pass, $0.038 per task at the listed price. Grok 4.5 used 6.9% of SuperGrok's weekly limit per pass, $0.032 per task at the listed price. Per task, Grok 4.5 is about 1.2x cheaper ($0.032 vs $0.038). Grok 4.5 is faster on real-world issues: a median task takes 3m 9s against 4m 53s. Both finished every task on it. The data gives no reason to pick Gemini 3.1 Pro over Grok 4.5 this week.

Changed since the previous graded week: Gemini 3.1 Pro: real-world issues usage 9.8% to 14.8% of the weekly limit, spec planning usage 0.9% to 1.1% of the weekly limit, vibe coding usage 1.0% to 0.7% of the weekly limit; Grok 4.5: real-world issues usage 8.0% to 6.9% of the weekly limit.

Gemini 3.1 ProClears tier 3 (partial tier 5) at 14.8% of Google AI Pro per runFOVERALLGrok 4.5Clears tier 4 (partial tier 5) at 6.9% of SuperGrok per runC-OVERALL
real-world@v3D+vsC-
DimensionGemini 3.1 ProGrok 4.5Advantage
Capability53%60%+7% (Grok 4.5 leads)
Median task time4m 53s3m 9s1m 44s faster (Grok 4.5 leads)
Approx. task cost$0.045 / task$0.032 / task$0.014 / task lower (Grok 4.5 leads)

Gemini 3.1 Pro: 14.8% of Google AI Pro (88.9% / 5h, 14.8% / 1w) per run. Its weekly limit holds 6 of the week's 33.6 5h windows. Grok 4.5: 6.9% of SuperGrok (6.9% / 1w) per run. Each figure is a share of that agent's own plan, so there is no delta to draw.

spec-planning@v1Fvs-

Only Gemini 3.1 Pro has run this suite, so there is no head-to-head yet.

vibe-coding@v1FvsC-
DimensionGemini 3.1 ProGrok 4.5Advantage
Capability27%59%+32% (Grok 4.5 leads)
Median task time4m 38s3m 40s58s faster (Grok 4.5 leads)
Approx. task cost$0.034 / task$0.031 / task$0.004 / task lower (Grok 4.5 leads)

Gemini 3.1 Pro: 0.7% of Google AI Pro (4.4% / 5h, 0.7% / 1w) per good solve. Its weekly limit holds 6 of the week's 33.6 5h windows. Grok 4.5: 0.4% of SuperGrok (0.4% / 1w) per good solve. Each figure is a share of that agent's own plan, so there is no delta to draw.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, Gemini 3.1 Pro or Grok 4.5?

Grok 4.5 clears higher: tier 4 of 5 versus tier 3 for Gemini 3.1 Pro on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-09-28.

Which is cheaper per task, Gemini 3.1 Pro or Grok 4.5?

Gemini 3.1 Pro used 15.9% of Google AI Pro's weekly limit per pass, $0.038 per task at the listed price. Grok 4.5 used 6.9% of SuperGrok's weekly limit per pass, $0.032 per task at the listed price. Per task, Grok 4.5 is about 1.2x cheaper ($0.032 vs $0.038).