Electricity Bench

Gemini 3.7 Flash vs GPT-5.6 Sol

Gemini 3.7 Flash was superseded by Gemini 3.8 Flash. GPT-5.6 Sol was superseded by GPT-6.1 Sol. The grades below are the last ones measured, so read them as history rather than as this week's board. See the current roster.

GPT-5.6 Sol leads overall, C- to D+. Gemini 3.7 Flash is about 1.6x cheaper per task. By suite, real-world issues is even at C- and GPT-5.6 Sol takes spec planning (C+ to D-). Take Gemini 3.7 Flash anyway when you want the side that is cheaper per task, faster on real-world issues and ahead on vibe coding.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: GPT-5.6 Sol vs Gemini 3.7 Flash

GPT-5.6 Sol leads overall, C- to D+. By suite: real-world issues is even at C- (60% vs 60% capability); GPT-5.6 Sol takes spec planning, C+ to D- (68% vs 40% capability); Gemini 3.7 Flash takes vibe coding, C- to D+ (61% vs 51% capability). Gemini 3.7 Flash used 9.4% of Google AI Pro's weekly limit per pass, $0.023 per task at the listed price. GPT-5.6 Sol used 3% of ChatGPT Pro 5x's weekly limit per pass, $0.036 per task at the listed price. Per task, Gemini 3.7 Flash is about 1.6x cheaper ($0.023 vs $0.036). Gemini 3.7 Flash is faster on real-world issues: a median task takes 2m 38s against 3m 13s. Both finished every task on it. Reasons to pick Gemini 3.7 Flash anyway: it is cheaper per task ($0.023 vs $0.036), it is faster on real-world issues (2m 38s vs 3m 13s median) and it leads on vibe coding (C- to D+).

Changed since the previous graded week: Gemini 3.7 Flash: real-world issues usage 8.5% to 8.2% of the weekly limit, spec planning D to D-, spec planning usage 1.3% to 1.2% of the weekly limit; GPT-5.6 Sol: vibe coding C- to D+.

Gemini 3.7 FlashClears tier 4 (partial tier 5) at 8.2% of Google AI Pro per runD+OVERALLGPT-5.6 SolClears tier 4 (partial tier 5) at 2.0% of ChatGPT Pro 5x per runC-OVERALL
real-world@v3C-vsC-
DimensionGemini 3.7 FlashGPT-5.6 SolAdvantage
Capability60%60%Even
Median task time2m 38s3m 13s35s faster (Gemini 3.7 Flash leads)
Approx. task cost$0.025 / task$0.031 / task$0.006 / task lower (Gemini 3.7 Flash leads)

Gemini 3.7 Flash: 8.2% of Google AI Pro (49.1% / 5h, 8.2% / 1w) per run. Its weekly limit holds 6 of the week's 33.6 5h windows. GPT-5.6 Sol: 2.0% of ChatGPT Pro 5x (2.0% / 1w) per run. Each figure is a share of that agent's own plan, so there is no delta to draw.

spec-planning@v1D-vsC+
DimensionGemini 3.7 FlashGPT-5.6 SolAdvantage
Capability40%68%+29% (GPT-5.6 Sol leads)
Median task time1m 38s6m 0s4m 22s faster (Gemini 3.7 Flash leads)
Approx. task cost$0.014 / task$0.057 / task$0.044 / task lower (Gemini 3.7 Flash leads)

Gemini 3.7 Flash: 1.2% of Google AI Pro (7.1% / 5h, 1.2% / 1w) per run. Its weekly limit holds 6 of the week's 33.6 5h windows. GPT-5.6 Sol: 1.0% of ChatGPT Pro 5x (1.0% / 1w) per run. Each figure is a share of that agent's own plan, so there is no delta to draw.

vibe-coding@v1C-vsD+
DimensionGemini 3.7 FlashGPT-5.6 SolAdvantage
Capability61%51%+10% (Gemini 3.7 Flash leads)
Median task time3m 37s3m 56s19s faster (Gemini 3.7 Flash leads)
Approx. task cost$0.028 / task$0.092 / task$0.064 / task lower (Gemini 3.7 Flash leads)

Gemini 3.7 Flash: 0.6% of Google AI Pro (3.6% / 5h, 0.6% / 1w) per good solve. Its weekly limit holds 6 of the week's 33.6 5h windows. GPT-5.6 Sol: 0.4% of ChatGPT Pro 5x (0.4% / 1w) per good solve. Each figure is a share of that agent's own plan, so there is no delta to draw.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, Gemini 3.7 Flash or GPT-5.6 Sol?

They are even on our ladder: both clear tier 4 of 5 at 60% capability on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-08-30.

Which is cheaper per task, Gemini 3.7 Flash or GPT-5.6 Sol?

Gemini 3.7 Flash used 9.4% of Google AI Pro's weekly limit per pass, $0.023 per task at the listed price. GPT-5.6 Sol used 3% of ChatGPT Pro 5x's weekly limit per pass, $0.036 per task at the listed price. Per task, Gemini 3.7 Flash is about 1.6x cheaper ($0.023 vs $0.036).