kimi-cli + kimi-k3-256k
kimi-cli harness · Kimi Allegretto ($39/mo)
Clears tier 2 of 5 (partial at tiers 3 and 5) on Real-world issues
Capability mean across 3 graded suite families.
What the tiers mean
"Clears" marks the highest rung where every task at it and below was solved. Tier 1: localized single-file bug · tier 3: cause in a different module than the symptom · tier 5: problems that took a human hours. A cleared rung solved every task on it; a partial rung solved some. The headline stops at the last unbroken rung, so a solved tier above a half-solved one counts only as partial.
What the levels mean
Level 1: a full brief naming the surface and the acceptance · level 3: the report as filed · level 5: a vague vibe report in a non-technical user's words. Context falls as the level rises, so vaguer is harder. A cleared level fixed every bug at it and below; a partial level fixed some. The headline stops at the last unbroken level, so a fix above a half-cleared level counts only as partial.
6 of this suite's tasks were credited rather than run: on a suite that asks the same problem at several levels of detail, solving it from the vaguest description credits the more detailed ones instead of asking again.
Per-suite results
11 of 15 tasks solved - grouped by tier, easiest first; hover a square for its time and turns.
Run on kimi-cli 0.36.1, model kimi-k3-256k (--thinking-effort low), week 2026-W33.
One run of this suite consumes 99.0% of Kimi Allegretto's 5h window and 20.0% of the weekly limit.
As work per window: about 148 suite runs per month on this plan - 3.78 runs per plan-dollar (at $39/mo, priced 2026-07-20). Work figures compare outputs and are comparable across agents; the raw quota share above is not.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Moonshot account on eb-runner, shared with the IDE usage panes on the claude host; quiet-window until a dedicated benchmark login exists.
0 of 4 tasks solved - hover a square for its time and turns.
Run on kimi-cli 0.36.1, model kimi-k3-256k (--thinking-effort low), week 2026-W34.
One run of this suite consumes 51.0% of Kimi Allegretto's 5h window and 10.0% of the weekly limit.
As work per window: about 286 suite runs per month on this plan - 7.35 runs per plan-dollar (at $39/mo, priced 2026-07-20). Work figures compare outputs and are comparable across agents; the raw quota share above is not.
How the plan cost was measured
Measured per suite run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Moonshot account on eb-runner, shared with the IDE usage panes on the claude host; quiet-window until a dedicated benchmark login exists.
16 of 25 tasks solved (6 credited without running) - each row is one bug, vaguest report first and context growing to the right; hover a square for its time and turns.
Run on kimi-cli 0.36.1, model kimi-k3-256k (--thinking-effort low), week 2026-W34.
One good solve on this suite consumes, on average, 10.2% of Kimi Allegretto's 5h window and 2.0% of the weekly limit.
As work per window: about 1432 good solves per month on this plan - 36.73 solves per plan-dollar (at $39/mo, priced 2026-07-20). Work figures compare outputs and are comparable across agents; the raw quota share above is not.
How the plan cost was measured
Counted over each issue's first solve (5 solves); the walk's failed harder cells and confirmation runs are the benchmark's own search cost and are excluded (full walk: 181.0% over 19 cells). Measured per cell run (reported-delta); an absolute cost on this plan, not comparable with another agent's plan. Measured in a declared quiet window on a shared subscription: operator Moonshot account on eb-runner, shared with the IDE usage panes on the claude host; quiet-window until a dedicated benchmark login exists.
Task outputs from these runs are withheld: publishing them would reveal the private suite. Per-task scores, times and turns are in the chips above.
Trajectory
One point per grading week - a re-run within a week shows only its latest execution. Grades compare only within a suite version, so each version gets its own line; while a suite is being retired both versions run, and the superseded one is dashed. The filled point is the run the scorecard shows today. Click a point or week for plan usage. All agents over time →
Weekly heads
Newest week first. Click a row for usage.
| Week | Grade | Capability | Usage |
|---|---|---|---|
| 2026-W34 · v3now | C | 57% | 99.0% |
| 2026-W33 · v3 | C | 59% | 156.0% |
Frequently asked questions
Is kimi-k3-256k worth it on Kimi Allegretto?
On our private real-world suite (real-world@v3), kimi-k3-256k earned grade C with 57% capability, clearing every problem up to tier 2 of 5 - a bug spanning a few files and landing some tier 5 problems. It runs on Kimi Allegretto at $39/month (public price as of 2026-07-20). One full suite run consumed 99.0% of the plan's 5h window. Measured 2026-08-16.
How good is kimi-k3-256k at real-world coding?
kimi-k3-256k clears tier 2 of our five-tier real-world ladder - every task at that tier and below - where tier 1 is a localized single-file bug and tier 5 is a problem that took a human engineer hours. It solves some but not all tasks up at tier 5. Its capability score on real-world@v3 is 57% (grade C). Last benchmarked 2026-08-16.
How much coding work does Kimi Allegretto buy with kimi-k3-256k?
Roughly 148 suite runs per month on Kimi Allegretto, or 3.78 runs per plan-dollar at $39/month (priced 2026-07-20). Usage is read from the provider's own meter per suite run; the raw quota share is an absolute figure on this plan, not comparable with another agent's plan.
Compare
- kimi-k3-256k vs gemini-3.1-pro
- kimi-k3-256k vs gemini-3.6-flash
- kimi-k3-256k vs gemini-3.7-flash
- kimi-k3-256k vs claude-fable-5
- kimi-k3-256k vs claude-haiku-4-5
- kimi-k3-256k vs claude-opus-5
- kimi-k3-256k vs claude-sonnet-5
- kimi-k3-256k vs gpt-5.6-luna
- kimi-k3-256k vs gpt-5.6-sol
- kimi-k3-256k vs gpt-5.6-terra
- kimi-k3-256k vs composer-2.5
- kimi-k3-256k vs grok-4.5
- kimi-k3-256k vs grok-4.6