Electricity Bench

Gemini 3.1 Pro vs Grok 4.6

Grok 4.6 was superseded by Grok 4.7. The grades below are the last ones measured, so read them as history rather than as this week's board. See the current roster.

Grok 4.6 leads overall, C- to F. Gemini 3.1 Pro is about 4.1x cheaper per task. By suite, Gemini 3.1 Pro takes real-world issues (D+ to D) and Grok 4.6 takes spec planning (C to F). Take Gemini 3.1 Pro anyway when you want the side that is cheaper per task, faster on real-world issues and ahead on real-world issues.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: Grok 4.6 vs Gemini 3.1 Pro

Grok 4.6 leads overall, C- to F. By suite: Gemini 3.1 Pro takes real-world issues, D+ to D (53% vs 47% capability); Grok 4.6 takes spec planning, C to F (66% vs 21% capability); Grok 4.6 takes vibe coding, C- to F (61% vs 27% capability). Gemini 3.1 Pro used 15.9% of Google AI Pro's weekly limit per pass, $0.038 per task at the listed price. Grok 4.6 used 13% of SuperGrok Plus's weekly limit per pass, $0.157 per task at the listed price. Per task, Gemini 3.1 Pro is about 4.1x cheaper ($0.038 vs $0.157). Gemini 3.1 Pro is faster on real-world issues: a median task takes 4m 53s against 7m 10s. Both finished every task on it. Reasons to pick Gemini 3.1 Pro anyway: it is cheaper per task ($0.038 vs $0.157), it is faster on real-world issues (4m 53s vs 7m 10s median) and it leads on real-world issues (D+ to D).

Changed since the previous graded week: Gemini 3.1 Pro: real-world issues usage 9.8% to 14.8% of the weekly limit, spec planning usage 0.9% to 1.1% of the weekly limit, vibe coding usage 1.0% to 0.7% of the weekly limit; Grok 4.6: real-world issues C- to D, real-world issues usage 4% to 3% of the weekly limit, vibe coding C+ to C-.

Gemini 3.1 ProClears tier 3 (partial tier 5) at 14.8% of Google AI Pro per runFOVERALLGrok 4.6Clears tier 4 (partial tier 5) at 3.0% of SuperGrok Plus per runC-OVERALL
real-world@v3D+vsD
DimensionGemini 3.1 ProGrok 4.6Advantage
Capability53%47%+7% (Gemini 3.1 Pro leads)
Median task time4m 53s7m 10s2m 17s faster (Gemini 3.1 Pro leads)
Approx. task cost$0.045 / task$0.046 / task$0.001 / task lower (Gemini 3.1 Pro leads)

Gemini 3.1 Pro: 14.8% of Google AI Pro (88.9% / 5h, 14.8% / 1w) per run. Its weekly limit holds 6 of the week's 33.6 5h windows. Grok 4.6: 3.0% of SuperGrok Plus (3.0% / 1w) per run. Each figure is a share of that agent's own plan, so there is no delta to draw.

spec-planning@v1FvsC
DimensionGemini 3.1 ProGrok 4.6Advantage
Capability21%66%+45% (Grok 4.6 leads)
Median task time1m 46s45m 26s43m 40s faster (Gemini 3.1 Pro leads)
Approx. task cost$0.013 / task$0.575 / task$0.562 / task lower (Gemini 3.1 Pro leads)

Gemini 3.1 Pro: 1.1% of Google AI Pro (6.5% / 5h, 1.1% / 1w) per run. Its weekly limit holds 6 of the week's 33.6 5h windows. Grok 4.6: 10.0% of SuperGrok Plus (10.0% / 1w) per run. Each figure is a share of that agent's own plan, so there is no delta to draw.

vibe-coding@v1FvsC-
DimensionGemini 3.1 ProGrok 4.6Advantage
Capability27%61%+34% (Grok 4.6 leads)
Median task time4m 38s11m 14s6m 36s faster (Gemini 3.1 Pro leads)
Approx. task cost$0.034 / task<$0.230 / task$0.196 / task lower (Gemini 3.1 Pro leads)

Gemini 3.1 Pro: 0.7% of Google AI Pro (4.4% / 5h, 0.7% / 1w) per good solve. Its weekly limit holds 6 of the week's 33.6 5h windows. Grok 4.6: under 1% of SuperGrok Plus per good solve. Each figure is a share of that agent's own plan, so there is no delta to draw.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, Gemini 3.1 Pro or Grok 4.6?

Grok 4.6 clears higher: tier 4 of 5 versus tier 3 for Gemini 3.1 Pro on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-09-28.

Which is cheaper per task, Gemini 3.1 Pro or Grok 4.6?

Gemini 3.1 Pro used 15.9% of Google AI Pro's weekly limit per pass, $0.038 per task at the listed price. Grok 4.6 used 13% of SuperGrok Plus's weekly limit per pass, $0.157 per task at the listed price. Per task, Gemini 3.1 Pro is about 4.1x cheaper ($0.038 vs $0.157).