Electricity Bench

Gemini 3.8 Flash vs GPT-6.1 Sol

GPT-6.1 Sol leads overall, C to C-. GPT-6.1 Sol is about 4.0x cheaper per task. By suite, real-world issues is even at C- and GPT-6.1 Sol takes spec planning (C+ to D+). The data gives no reason to pick Gemini 3.8 Flash over GPT-6.1 Sol this week.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: GPT-6.1 Sol vs Gemini 3.8 Flash

GPT-6.1 Sol leads overall, C to C-. By suite: real-world issues is even at C- (60% vs 57% capability); GPT-6.1 Sol takes spec planning, C+ to D+ (73% vs 50% capability); GPT-6.1 Sol takes vibe coding, C+ to C- (72% vs 56% capability). Gemini 3.8 Flash used 23.7% of Google AI Pro's weekly limit per pass, $0.057 per task at the listed price. GPT-6.1 Sol used 6% of ChatGPT Plus's weekly limit per pass, $0.015 per task at the listed price; at the ~$100 tier the board defaults to that is an estimated ~1.2% of ChatGPT Pro 5x's weekly limit (est, about $0.015 per task). Per task, GPT-6.1 Sol is about 4.0x cheaper ($0.015 vs $0.057). GPT-6.1 Sol is faster on real-world issues: a median task takes 3m 26s against 5m 21s. Both finished every task on it. The data gives no reason to pick Gemini 3.8 Flash over GPT-6.1 Sol this week.

Changed since W39: Gemini 3.8 Flash: real-world issues F to C-, real-world issues usage 18.0% to 21.9% of the weekly limit, spec planning usage 2.0% to 1.8% of the weekly limit, vibe coding D+ to C-, vibe coding usage 1.2% to 1.1% of the weekly limit; GPT-6.1 Sol has one graded week under this version (W40).

Gemini 3.8 FlashClears tier 4 (partial tier 5) at 21.9% of Google AI Pro per runC-OVERALLGPT-6.1 SolClears tier 2 (partial tier 5) at 4.0% of ChatGPT Plus per runCOVERALL
real-world@v3C-vsC-
DimensionGemini 3.8 FlashGPT-6.1 SolAdvantage
Capability60%57%+3% (Gemini 3.8 Flash leads)
Median task time5m 21s3m 26s1m 55s faster (GPT-6.1 Sol leads)
Approx. task cost$0.067 / task$0.012 / task$0.055 / task lower (GPT-6.1 Sol leads)

Gemini 3.8 Flash: 21.9% of Google AI Pro (131.3% / 5h, 21.9% / 1w) per run. Its weekly limit holds 6 of the week's 33.6 5h windows. GPT-6.1 Sol: 4.0% of ChatGPT Plus (28.0% / 5h, 4.0% / 1w) per run. Its weekly limit holds about 6 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

spec-planning@v1D+vsC+
DimensionGemini 3.8 FlashGPT-6.1 SolAdvantage
Capability50%73%+23% (GPT-6.1 Sol leads)
Median task time3m 0s6m 22s3m 22s faster (Gemini 3.8 Flash leads)
Approx. task cost$0.021 / task$0.023 / task$0.002 / task lower (Gemini 3.8 Flash leads)

Gemini 3.8 Flash: 1.8% of Google AI Pro (11.0% / 5h, 1.8% / 1w) per run. Its weekly limit holds 6 of the week's 33.6 5h windows. GPT-6.1 Sol: 2.0% of ChatGPT Plus (9.0% / 5h, 2.0% / 1w) per run. Its weekly limit holds about 6 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

vibe-coding@v1C-vsC+
DimensionGemini 3.8 FlashGPT-6.1 SolAdvantage
Capability56%72%+15% (GPT-6.1 Sol leads)
Median task time7m 51s4m 55s2m 56s faster (GPT-6.1 Sol leads)
Approx. task cost$0.048 / task$0.018 / task$0.030 / task lower (GPT-6.1 Sol leads)

Gemini 3.8 Flash: 1.1% of Google AI Pro (6.3% / 5h, 1.1% / 1w) per good solve. Its weekly limit holds 6 of the week's 33.6 5h windows. GPT-6.1 Sol: 0.4% of ChatGPT Plus (2.2% / 5h, 0.4% / 1w) per good solve. Its weekly limit holds about 6 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, Gemini 3.8 Flash or GPT-6.1 Sol?

Gemini 3.8 Flash clears higher: tier 4 of 5 versus tier 2 for GPT-6.1 Sol on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-09-27.

Which is cheaper per task, Gemini 3.8 Flash or GPT-6.1 Sol?

Gemini 3.8 Flash used 23.7% of Google AI Pro's weekly limit per pass, $0.057 per task at the listed price. GPT-6.1 Sol used 6% of ChatGPT Plus's weekly limit per pass, $0.015 per task at the listed price; at the ~$100 tier the board defaults to that is an estimated ~1.2% of ChatGPT Pro 5x's weekly limit (est, about $0.015 per task). Per task, GPT-6.1 Sol is about 4.0x cheaper ($0.015 vs $0.057).