Electricity Bench

Composer 2.5 vs Kimi K3 256K

Kimi K3 256K leads overall, D+ to D. Composer 2.5 is about 6.8x cheaper per task. By suite, real-world issues is even at C- and Kimi K3 256K takes spec planning (C- to D). Take Composer 2.5 anyway when you want the side that is cheaper per task and faster on real-world issues.Measured

Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.

Verdict: Kimi K3 256K vs Composer 2.5

Kimi K3 256K leads overall, D+ to D. By suite: real-world issues is even at C- (57% vs 60% capability); Kimi K3 256K takes spec planning, C- to D (56% vs 48% capability); vibe coding is even at F (23% vs 34% capability). Composer 2.5 used 2.0% of Cursor Pro's monthly limit per pass, $0.021 per task at the listed price. Kimi K3 256K used 31% of Kimi Allegretto's weekly limit per pass, $0.146 per task at the listed price. Per task, Composer 2.5 is about 6.8x cheaper ($0.021 vs $0.146). Composer 2.5 is faster on real-world issues: a median task takes 1m 34s against 3m 31s. Both finished every task on it. Reasons to pick Composer 2.5 anyway: it is cheaper per task ($0.021 vs $0.146) and it is faster on real-world issues (1m 34s vs 3m 31s median).

Changed since W39: Composer 2.5: real-world issues D+ to C-; Kimi K3 256K: real-world issues usage 23% to 21% of the weekly limit, spec planning D+ to C-, spec planning usage 11% to 10% of the weekly limit, vibe coding usage 3% to 2.6% of the weekly limit.

Composer 2.5Clears tier 1 (partial tier 5) at 1.6% of Cursor Pro per runDOVERALLKimi K3 256KClears tier 4 (partial tier 5) at 104.0% of Kimi Allegretto per runD+OVERALL
real-world@v3C-vsC-
DimensionComposer 2.5Kimi K3 256KAdvantage
Capability57%60%+3% (Kimi K3 256K leads)
Median task time1m 34s3m 31s1m 57s faster (Composer 2.5 leads)
Approx. task cost$0.021 / task$0.126 / task$0.104 / task lower (Composer 2.5 leads)

Composer 2.5: 1.6% of Cursor Pro (1.6% / 1m) per run. Kimi K3 256K: 104.0% of Kimi Allegretto (104.0% / 5h, 21.0% / 1w) per run. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

Work per window: Composer 2.5 63 vs Kimi K3 256K 140 suite runs per month (Kimi K3 256K leads); per plan-dollar: 3.15 vs 3.60 (Kimi K3 256K leads). Work figures compare outputs, so unlike quota shares they are comparable across plans.

spec-planning@v1DvsC-
DimensionComposer 2.5Kimi K3 256KAdvantage
Capability48%56%+8% (Kimi K3 256K leads)
Median task time1m 8s5m 26s4m 18s faster (Composer 2.5 leads)
Approx. task cost$0.022 / task$0.224 / task$0.202 / task lower (Composer 2.5 leads)

Composer 2.5: 0.4% of Cursor Pro (0.4% / 1m) per run. Kimi K3 256K: 52.0% of Kimi Allegretto (52.0% / 5h, 10.0% / 1w) per run. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

Work per window: Composer 2.5 225 vs Kimi K3 256K 281 suite runs per month (Kimi K3 256K leads); per plan-dollar: 11.23 vs 7.20 (Composer 2.5 leads). Work figures compare outputs, so unlike quota shares they are comparable across plans.

vibe-coding@v1FvsF
DimensionComposer 2.5Kimi K3 256KAdvantage
Capability23%34%+12% (Kimi K3 256K leads)
Median task time1m 37s7m 33s5m 56s faster (Composer 2.5 leads)
Approx. task cost$0.022 / task$0.233 / task$0.211 / task lower (Composer 2.5 leads)

Composer 2.5: 0.1% of Cursor Pro (0.1% / 1m) per good solve. Kimi K3 256K: 13.2% of Kimi Allegretto (13.2% / 5h, 2.6% / 1w) per good solve. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.

Work per window: Composer 2.5 920 vs Kimi K3 256K 1107 good solves per month (Kimi K3 256K leads); per plan-dollar: 46.02 vs 28.38 (Composer 2.5 leads). Work figures compare outputs, so unlike quota shares they are comparable across plans.

Grades and deltas are comparable only within a suite version. See methodology.

Frequently asked questions

Which is better for coding, Composer 2.5 or Kimi K3 256K?

Kimi K3 256K clears higher: tier 4 of 5 versus tier 1 for Composer 2.5 on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-09-27.

Which is cheaper per task, Composer 2.5 or Kimi K3 256K?

Composer 2.5 used 2.0% of Cursor Pro's monthly limit per pass, $0.021 per task at the listed price. Kimi K3 256K used 31% of Kimi Allegretto's weekly limit per pass, $0.146 per task at the listed price. Per task, Composer 2.5 is about 6.8x cheaper ($0.021 vs $0.146).

Which gives more coding work per dollar, Composer 2.5 or Kimi K3 256K?

Kimi K3 256K delivers more benchmark work per subscription dollar: 3.60 suite runs per plan-dollar versus 3.15 for Composer 2.5 (at $39/month priced 2026-07-20, and $20/month priced 2026-07-20). Work figures compare outputs, so unlike raw quota shares they are comparable across plans.

What do Composer 2.5 and Kimi K3 256K cost a month?

Composer 2.5 runs on Cursor Pro at $20/month (public price as of 2026-07-20). Kimi K3 256K runs on Kimi Allegretto at $39/month (public price as of 2026-07-20).