Electricity Bench

Gemini 3.1 Pro

antigravity-cli harness · Google AI Pro

Coding agent review, updated weekly (latest W40 2026).

Gemini 3.1 Pro scores F overall on Electricity Bench and clears tier 3 of 5 on real-world issues. Its best suite is spec planning at F. One pass of our task set uses 15.9% of Google AI Pro's weekly limit. In our own use it has no clear role right now.Measured

FOVERALL

Clears tier 3 of 5 (partial at tiers 4 and 5) on Real-world issues

Capability mean across 3 graded suite families.

capability
53%
plan cost per run
14.8% of Google AI Pro
median task time
4m 53s
reliability
100%
12345
What the tiers mean

"Clears" marks the highest rung where every task at it and below was solved. Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some. The headline stops at the last unbroken rung, so a solved tier above a half-solved one counts only as partial.

Clears no level (partial at levels 1, 2, 3, 4 and 5) on Vibe coding v1 - the vaguest report level at which it still fixed every bug, having also cleared every richer level below
12345
What the levels mean

Level 1: a full brief naming the surface and the acceptance · level 3: the report as filed · level 5: a vague vibe report in a non-technical user's words. Context falls as the level rises, so vaguer is harder. A cleared level fixed every bug at it and below; a partial level fixed some. The headline stops at the last unbroken level, so a fix above a half-cleared level counts only as partial.

3 of this suite's tasks were credited rather than run: on a suite that asks the same problem at several levels of detail, solving it from the vaguest description credits the more detailed ones instead of asking again.

Per-suite results

D+
Real-world issues v3
53%
plan cost per run
14.8% of Google AI Pro
median task time
4m 53s
reliability
100%
tier 1tier 2tier 3tier 4tier 5

11 of 15 tasks solved - grouped by tier, easiest first; hover a square for its time and turns.

Run on antigravity-cli 1.2.12, model gemini-3.1-pro (effort high), week 2026-W40.

One run of this suite consumes 88.9% of Google AI Pro's 5h window and 14.8% of the weekly limit.

Google AI Pro's weekly limit holds 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Measured over 653 runs in 41 session windows, every one of which agreed.

How the plan cost was measured

Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a dedicated subscription.

How the limits work on Antigravity CLI: /limits/antigravity-cli/

F
Spec planning v1
21%
plan cost per run
1.1% of Google AI Pro
median task time
1m 46s
reliability
100%

0 of 4 tasks solved - hover a square for its time and turns.

Run on antigravity-cli 1.2.12, model gemini-3.1-pro (effort high), week 2026-W40.

One run of this suite consumes 6.5% of Google AI Pro's 5h window and 1.1% of the weekly limit.

Google AI Pro's weekly limit holds 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Measured over 653 runs in 41 session windows, every one of which agreed.

How the plan cost was measured

Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a dedicated subscription.

How the limits work on Antigravity CLI: /limits/antigravity-cli/

F
Vibe coding v1
27%
plan cost per solve
0.7% of Google AI Pro
median task time
4m 38s
reliability
100%
level 5level 4level 3level 2level 1issue aissue bissue cissue dissue e

11 of 25 tasks solved (3 credited without running) - each row is one bug, vaguest report first and context growing to the right; hover a square for its time and turns.

Run on antigravity-cli 1.2.12, model gemini-3.1-pro (effort high), week 2026-W40.

One good solve on this suite consumes, on average, 4.4% of Google AI Pro's 5h window and 0.7% of the weekly limit.

Google AI Pro's weekly limit holds 6 fully spent 5h windows, out of the 33.6 a calendar week contains. Measured over 653 runs in 41 session windows, every one of which agreed.

How the plan cost was measured

Counted over each issue's first solve (4 solves); the walk's failed harder cells and confirmation runs are the benchmark's own search cost and are excluded (full walk: 17.2% over 22 cells). Measured per cell run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a dedicated subscription.

How the limits work on Antigravity CLI: /limits/antigravity-cli/

Task outputs from these runs are withheld: publishing them would reveal the private suite. Per-task scores, times and turns are in the chips above.

Trajectory

One point per grading week - a re-run within a week shows only its latest execution. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. Click a point or week for plan usage. All agents over time →

Suite
Real-world issues · 9 weeks graded · latest 2026-W40
20%30%40%50%60%70%80%2026-W322026-W342026-W362026-W382026-W40Gemini 3.1 Pro · 2026-W32 · Aug 5, 2026 · 47% · DGemini 3.1 Pro · 2026-W33 · Aug 10, 2026 · 47% · DGemini 3.1 Pro · 2026-W34 · Aug 17, 2026 · 53% · D+Gemini 3.1 Pro · 2026-W35 · Aug 23, 2026 · 67% · CGemini 3.1 Pro · 2026-W36 · Aug 31, 2026 · 47% · DGemini 3.1 Pro · 2026-W37 · Sep 6, 2026 · 47% · DGemini 3.1 Pro · 2026-W38 · Sep 14, 2026 · 33% · FGemini 3.1 Pro · 2026-W39 · Sep 21, 2026 · 53% · D+Gemini 3.1 Pro · 2026-W40 · Sep 28, 2026 · 53% · D+

Weekly heads

Newest week first. Click a row for usage.

WeekGradeCapabilityClearsTask timeUsage /1w
2026-W40 · v3nowD+53%tier 3 · p54m 53s14.8% of Google AI Pro
2026-W39 · v3D+53%tier 3 · p52m 55s9.8% of Google AI Pro
2026-W38 · v3F33%tier 3 · p53m 30s11.2% of Google AI Pro
2026-W37 · v3D47%tier 3 · p53m 27s12.4% of Google AI Pro
2026-W36 · v3D47%tier 3 · p53m 14s9.0% of Google AI Pro
2026-W35 · v3C67%tier 3 · p53m 34s9.5% of Google AI Pro
2026-W34 · v3D+53%tier 3 · p53m 25s8.5% of Google AI Pro
2026-W33 · v3D47%tier 3 · p54m 24s8.0% of Google AI Pro
2026-W32 · v3D47%tier 4 · p52m 31s16.8% of Google AI Pro

Embed this grade

Put Gemini 3.1 Pro's current grade in a README or a docs page.

Gemini 3.1 Pro on Electricity Bench: F

Markdown
[![Gemini 3.1 Pro on Electricity Bench: F](https://electricitybench.com/badge/antigravity-cli-gemini-3-1-pro.svg)](https://electricitybench.com/models/antigravity-cli-gemini-3-1-pro/)
HTML
<a href="https://electricitybench.com/models/antigravity-cli-gemini-3-1-pro/"><img src="https://electricitybench.com/badge/antigravity-cli-gemini-3-1-pro.svg" alt="Gemini 3.1 Pro on Electricity Bench: F"></a>

Badges update every Monday with the weekly run. GitHub renders them from this URL.

Frequently asked questions

Is Gemini 3.1 Pro worth it on Google AI Pro?

On our private real-world suite (real-world@v3), Gemini 3.1 Pro earned grade D+ with 53% capability, clearing every problem up to tier 3 of 5 - a problem whose cause sits in a different module than the symptom and landing some tier 5 problems. It runs on Google AI Pro. One full suite run consumed 14.8% of the plan's week window. Measured 2026-09-28.

Is Gemini 3.1 Pro good for coding?

Gemini 3.1 Pro clears tier 3 of our five-tier real-world ladder - every task at that tier and below - where tier 1 is a localized single-file bug and tier 5 is a problem that took a human engineer hours. It solves some but not all tasks up at tier 5. Its capability score on real-world@v3 is 53% (grade D+). Last benchmarked 2026-09-28.

How much of Google AI Pro does one Gemini 3.1 Pro task use?

One full suite run consumed 14.8% of Google AI Pro's weekly limit, read off the provider's own meter before and after the run. A benchmark task is a whole multi-turn agent session, not a message, and the share belongs to this plan alone: it is not comparable with another agent's plan. Measured 2026-09-28.

Compare