Electricity Bench

antigravity-cli + gemini-3.6-flash

antigravity-cli harness · Google AI Pro

plan
Google AI Pro
price
no public price recorded
allowance
100 provider-percent per week
isolation
dedicated-account
usage method
reported-delta

Dedicated benchmark Google account ebbench.runner@gmail.com on eb-runner; no interactive or IDE use shares the AI Pro weekly Gemini window.

Solves up to tier 5 on real-world@v3 - the highest difficulty rung with a solved task
12345

Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some.

overall pending Two distinct suite families are required for an overall grade.

Per-suite results

C
real-world@v3
60%
plan cost per run
19.6% of Google AI Pro
latency med / p95
151841 ms / 401980 ms
reliability
100%
error rate
0%

Run on antigravity-cli 1.1.10, model gemini-3.6-flash (--effort medium), Aug 5, 2026.

One run of this suite consumes 19.6% of Google AI Pro - about 5.1 runs per week. Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured on a benchmark-only account (Dedicated benchmark Google account ebbench.runner@gmail.com on eb-runner; no interactive or IDE use shares the AI Pro weekly Gemini window.).

Trajectory

Every scored run of this agent, oldest first. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. All agents over time →

real-world · 2 runs · latest Aug 5, 2026
Capability
0%50%100%Aug 5antigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 60% · grade Cantigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 100% · grade A0%50%100%Aug 5antigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 60% · grade Cantigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 100% · grade A
Median latency
89732 ms124837 ms159942 msAug 5antigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 151841 ms · grade Cantigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 97833 ms · grade A89732 ms124837 ms159942 msAug 5antigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 151841 ms · grade Cantigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 97833 ms · grade A
Plan cost per run - a share of this agent's own plan, not a cross-agent number
0.0%11.2%22.5%Aug 5antigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 19.6% · grade Cantigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 0.9% · grade A0.0%11.2%22.5%Aug 5antigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 19.6% · grade Cantigravity-cli + gemini-3.6-flash · v3 · Aug 5, 2026 · 0.9% · grade A

Aug 5, 2026: C on v3 · Aug 5, 2026: A on v3

Compare

Regrade changelog

Transcript excerpts

Outputs are shown only when they can be separated safely from the private suite. Scores and run metadata remain visible when an output is withheld.

real-world@v3
Task#TurnsLatencyOutput
b2r5h81-361033 mswithheldPrivate-suite content
c8g2m41-62744 mswithheldPrivate-suite content
d9b4x61-151841 mswithheldPrivate-suite content
h5n2v71-482849 mswithheldPrivate-suite content
h7w3j51-167764 mswithheldPrivate-suite content
k4f9t71-127363 mswithheldPrivate-suite content
m6t3q81-70966 mswithheldPrivate-suite content
p1x9k41-367323 mswithheldPrivate-suite content
q8k3j61-338468 mswithheldPrivate-suite content
r3v8m51-249893 mswithheldPrivate-suite content
s2j7f41-211589 mswithheldPrivate-suite content
t5s8n21-54602 mswithheldPrivate-suite content
w4j7q21-139714 mswithheldPrivate-suite content
y3p7k11-83580 mswithheldPrivate-suite content
z6q1v91-79360 mswithheldPrivate-suite content