Electricity Bench

GPT-6 Luna vs Composer 2.5

GPT-6 Luna leads overall, D to D-. GPT-6 Luna is about 9.3x cheaper per task. By suite, real-world issues is even at D+ and GPT-6 Luna takes spec planning (C- to D). The data gives no reason to pick Composer 2.5 over GPT-6 Luna this week.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: Composer 2.5 vs GPT-6 Luna

GPT-6 Luna leads overall, D to D-. By suite: real-world issues is even at D+ (50% vs 50% capability); GPT-6 Luna takes spec planning, C- to D (61% vs 45% capability); vibe coding is even at F (33% vs 23% capability). GPT-6 Luna used 1% of ChatGPT Plus's weekly limit per pass, $0.002 per task at the listed price; at the ~$100 tier the board defaults to that is an estimated ~0.2% of ChatGPT Pro 5x's weekly limit (est, about $0.002 per task). Composer 2.5 used 2.1% of Cursor Pro's monthly limit per pass, $0.023 per task at the listed price. Per task, GPT-6 Luna is about 9.3x cheaper ($0.002 vs $0.023). GPT-6 Luna is faster on real-world issues: a median task takes 1m 0s against 1m 21s. Both finished every task on it. The data gives no reason to pick Composer 2.5 over GPT-6 Luna this week.

Changed since W38: Composer 2.5: spec planning D- to D.

GPT-6 LunaClears tier 2 (partial tier 5) at 1.0% of ChatGPT Plus per runDOVERALLComposer 2.5Clears tier 2 (partial tier 5) at 1.7% of Cursor Pro per runD-OVERALL
real-world@v3D+vsD+
DimensionGPT-6 LunaComposer 2.5Advantage
Capability50%50%Even
Median task time1m 0s1m 21s21s faster (GPT-6 Luna leads)
Approx. task cost$0.003 / task$0.023 / task$0.020 / task lower (GPT-6 Luna leads)

GPT-6 Luna: 1.0% of ChatGPT Plus (1.0% / 5h, 1.0% / 1w) per run. Its weekly limit holds about 6 of the week's 33.6 5h windows. Composer 2.5: 1.7% of Cursor Pro (1.7% / 1m) per run. Each figure is a share of that agent's own plan, so there is no delta to draw.

spec-planning@v1C-vsD
DimensionGPT-6 LunaComposer 2.5Advantage
Capability61%45%+15% (GPT-6 Luna leads)
Median task time1m 8s1m 50s43s faster (GPT-6 Luna leads)
Approx. task cost<$0.011 / task$0.022 / task$0.011 / task lower (GPT-6 Luna leads)

GPT-6 Luna: under 1% of ChatGPT Plus per run. Its weekly limit holds about 6 of the week's 33.6 5h windows. Composer 2.5: 0.4% of Cursor Pro (0.4% / 1m) per run. Each figure is a share of that agent's own plan, so there is no delta to draw.

vibe-coding@v1FvsF
DimensionGPT-6 LunaComposer 2.5Advantage
Capability33%23%+10% (GPT-6 Luna leads)
Median task time1m 37s1m 55s18s faster (GPT-6 Luna leads)
Approx. task costnot measured$0.031 / task-

GPT-6 Luna: under 1% of ChatGPT Plus (0.4% / 5h) per good solve. Its weekly limit holds about 6 of the week's 33.6 5h windows. Composer 2.5: 0.2% of Cursor Pro (0.2% / 1m) per good solve. Each figure is a share of that agent's own plan, so there is no delta to draw.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, GPT-6 Luna or Composer 2.5?

They are even on our ladder: both clear tier 2 of 5 at 50% capability on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-09-23.

Which is cheaper per task, GPT-6 Luna or Composer 2.5?

GPT-6 Luna used 1% of ChatGPT Plus's weekly limit per pass, $0.002 per task at the listed price; at the ~$100 tier the board defaults to that is an estimated ~0.2% of ChatGPT Pro 5x's weekly limit (est, about $0.002 per task). Composer 2.5 used 2.1% of Cursor Pro's monthly limit per pass, $0.023 per task at the listed price. Per task, GPT-6 Luna is about 9.3x cheaper ($0.002 vs $0.023).

What do GPT-6 Luna and Composer 2.5 cost a month?

GPT-6 Luna runs on ChatGPT Plus at $20/month (public price as of 2026-07-20). Composer 2.5 runs on Cursor Pro at $20/month (public price as of 2026-07-20).