GPT-5.6 Luna vs Kimi K3 256K
GPT-5.6 Luna was superseded by GPT-6 Luna. The grades below are the last ones measured, so read them as history rather than as this week's board. See the current roster.
Kimi K3 256K leads overall, D+ to D. GPT-5.6 Luna is the cheaper side per task (at most $0.012). By suite, Kimi K3 256K takes real-world issues (C- to D) and spec planning is even at C-. Take GPT-5.6 Luna anyway when you want the side that is cheaper per task and faster on real-world issues.Measured
Each agent graded on the same private suites, running headless through its own tools. Advantages are oriented so a win is a win, whichever way the metric runs.
Verdict: Kimi K3 256K vs GPT-5.6 Luna
Kimi K3 256K leads overall, D+ to D. By suite: Kimi K3 256K takes real-world issues, C- to D (60% vs 45% capability); spec planning is even at C- (57% vs 56% capability); vibe coding is even at F (25% vs 34% capability). GPT-5.6 Luna used <1% of ChatGPT Pro 5x's weekly limit per pass, at most $0.012 per task at the listed price. Kimi K3 256K used 31% of Kimi Allegretto's weekly limit per pass, $0.146 per task at the listed price. Per task, GPT-5.6 Luna costs at most $0.012 against Kimi K3 256K's $0.146, so GPT-5.6 Luna is the cheaper side. GPT-5.6 Luna is faster on real-world issues: a median task takes 3m 7s against 3m 31s. Both finished every task on it. Reasons to pick GPT-5.6 Luna anyway: it is cheaper per task (at most $0.012 vs $0.146) and it is faster on real-world issues (3m 7s vs 3m 31s median).
Changed since the previous graded week: GPT-5.6 Luna: spec planning D+ to C-, vibe coding C- to F; Kimi K3 256K: real-world issues usage 23% to 21% of the weekly limit, spec planning D+ to C-, spec planning usage 11% to 10% of the weekly limit, vibe coding usage 3% to 2.6% of the weekly limit.
| Dimension | GPT-5.6 Luna | Kimi K3 256K | Advantage |
|---|---|---|---|
| Capability | 45% | 60% | +15% (Kimi K3 256K leads) |
| Median task time | 3m 7s | 3m 31s | 24s faster (GPT-5.6 Luna leads) |
| Approx. task cost | <$0.015 / task | $0.126 / task | $0.110 / task lower (GPT-5.6 Luna leads) |
GPT-5.6 Luna: under 1% of ChatGPT Pro 5x per run. Kimi K3 256K: 104.0% of Kimi Allegretto (104.0% / 5h, 21.0% / 1w) per run. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.
| Dimension | GPT-5.6 Luna | Kimi K3 256K | Advantage |
|---|---|---|---|
| Capability | 57% | 56% | +1% (GPT-5.6 Luna leads) |
| Median task time | 3m 25s | 5m 26s | 2m 0s faster (GPT-5.6 Luna leads) |
| Approx. task cost | <$0.057 / task | $0.224 / task | $0.167 / task lower (GPT-5.6 Luna leads) |
GPT-5.6 Luna: under 1% of ChatGPT Pro 5x per run. Kimi K3 256K: 52.0% of Kimi Allegretto (52.0% / 5h, 10.0% / 1w) per run. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.
| Dimension | GPT-5.6 Luna | Kimi K3 256K | Advantage |
|---|---|---|---|
| Capability | 25% | 34% | +9% (Kimi K3 256K leads) |
| Median task time | 3m 40s | 7m 33s | 3m 52s faster (GPT-5.6 Luna leads) |
| Approx. task cost | <$0.230 / task | $0.233 / task | $0.003 / task lower (GPT-5.6 Luna leads) |
GPT-5.6 Luna: under 1% of ChatGPT Pro 5x per good solve. Kimi K3 256K: 13.2% of Kimi Allegretto (13.2% / 5h, 2.6% / 1w) per good solve. Its weekly limit holds about 5 of the week's 33.6 5h windows. Each figure is a share of that agent's own plan, so there is no delta to draw.
Grades and deltas are comparable only within a suite version. See methodology.
Frequently asked questions
Which is better for coding, GPT-5.6 Luna or Kimi K3 256K?
Kimi K3 256K clears higher: tier 4 of 5 versus tier 1 for GPT-5.6 Luna on real-world@v3, our private suite of real-world coding problems. Latest measurement 2026-09-21.
Which is cheaper per task, GPT-5.6 Luna or Kimi K3 256K?
GPT-5.6 Luna used <1% of ChatGPT Pro 5x's weekly limit per pass, at most $0.012 per task at the listed price. Kimi K3 256K used 31% of Kimi Allegretto's weekly limit per pass, $0.146 per task at the listed price. Per task, GPT-5.6 Luna costs at most $0.012 against Kimi K3 256K's $0.146, so GPT-5.6 Luna is the cheaper side.
What do GPT-5.6 Luna and Kimi K3 256K cost a month?
GPT-5.6 Luna runs on ChatGPT Pro 5x at $100/month (public price as of 2026-09-05). Kimi K3 256K runs on Kimi Allegretto at $39/month (public price as of 2026-07-20).