Kimi API Pricing 2026: K3, K2.7 Code, and K2.6 Compared
Moonshot AI now offers 3 models โ from the K3 flagship at $2.90/$14.00 to budget options at $0.95/$4.00 per 1M tokens.
Moonshot AI's Kimi lineup has expanded significantly. The new Kimi K3 is a flagship model with 1M token context and always-on reasoning, designed for long-form coding and knowledge work. It joins the existing budget models K2.7 Code and K2.6, giving developers three distinct price-performance points.
All Kimi Models: Full Pricing Comparison
| Model | Input / 1M | Output / 1M | Context | Tier |
|---|---|---|---|---|
| Kimi K3 | $2.90 | $14.00 | 1M | Mid |
| Kimi K2.7 Code | $0.95 | $4.00 | 256K | Budget |
| Kimi K2.6 | $0.95 | $4.00 | 256K | Budget |
Key detail: Kimi K3 supports automatic context caching with a reported ~90% cache hit rate. On cache hits, input cost drops to ยฅ2/1M (~$0.29) โ making effective input cost 90% cheaper than list price for workloads with repeated context like multi-turn conversations or RAG pipelines.
Kimi K3: The Flagship
Kimi K3 is Moonshot's answer to Claude Opus 5 and GPT-5.4. It's designed for "long-form programming and end-to-end knowledge work" with always-on reasoning (configurable intensity: low, high, max).
Kimi K3 specs:
- $2.90 input / $14.00 output per 1M tokens (list price)
- ~$0.29 input on cache hits (90% cache hit rate)
- 1M token context window โ same as Claude Opus 5
- Always-on reasoning with configurable intensity
- Tool calls, JSON mode, structured output supported
At $14.00/1M output, K3 is significantly more expensive on output than competitors like Claude Opus 5 ($25.00) or GPT-5.4 ($15.00). But the cache-hit pricing makes it competitive for input-heavy workloads โ a 10:1 input-to-output ratio with cached context brings effective cost well below list price.
How Kimi Compares to Other Providers
| Model | Input / 1M | Output / 1M | Context | Provider |
|---|---|---|---|---|
| Kimi K3 | $2.90 | $14.00 | 1M | Moonshot |
| DeepSeek V4 Pro ($0.66/$1.98) | 128K | DeepSeek | ||
| Claude Sonnet 5 | $2.00 | $10.00 | 200K | Anthropic |
| GPT-5.4 | $2.50 | $15.00 | 128K | OpenAI |
| Claude Opus 5 | $5.00 | $25.00 | 1M | Anthropic |
| Gemini 3.1 Pro | $2.00 | $12.00 | 1M |
Kimi K3 sits in the mid-tier on list price. Its main competitors for the "flagship with long context" use case are Claude Opus 5 ($5/$25, 1M) and Gemini 3.1 Pro ($2/$12, 1M). DeepSeek V4 Pro is dramatically cheaper but has only 128K context.
Budget Kimi Models: K2.7 Code and K2.6
| Model | Input / 1M | Output / 1M | Context | Best For |
|---|---|---|---|---|
| Kimi K2.7 Code | $0.95 | $4.00 | 256K | Code generation, technical tasks |
| Kimi K2.6 | $0.95 | $4.00 | 256K | General purpose, Chinese language |
| DeepSeek V4 Flash | $0.22 | $0.66 | 128K | Cheapest option |
| Claude Haiku 4.5 | $1.00 | $5.00 | 200K | Best budget quality |
| GPT-5.4 nano | $0.20 | $1.25 | 128K | OpenAI budget |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 1M | Cheapest with long context |
At $0.95/$4.00, the budget Kimi models are 6-7x more expensive than DeepSeek V4 Flash on input. They're comparable to Claude Haiku 4.5 ($1.00/$5.00) โ slightly cheaper on input, slightly cheaper on output. The main draw is Chinese language optimization and 256K context.
Real-World Cost Scenarios
| Scenario | Kimi K3 | Kimi K2.6 | Claude Sonnet 5 | DeepSeek V4 Pro |
|---|---|---|---|---|
| Chatbot (10K req/mo, 500 in / 1K out) | $154.50 | $47.50 | $110.00 | $11.15 |
| Code Gen (5K req/mo, 2K in / 3K out) | $239.00 | $69.50 | $170.00 | $17.95 |
| Doc Analysis (2K req/mo, 10K in / 1K out) | $86.00 | $27.00 | $60.00 | $11.54 |
| RAG Pipeline (20K req/mo, 2K in / 500 out) | $256.00 | $78.00 | $130.00 | $28.30 |
With K3 cache hits (90% rate): The RAG pipeline scenario drops from $256 to ~$68 โ making K3 competitive with Claude Sonnet 5 for cache-heavy workloads. Cache effectiveness is the key differentiator for K3 economics.
Which Kimi Model Should You Pick?
Decision guide:
- Complex coding, agentic work, or knowledge work with long documents? โ Kimi K3 (1M context, always-on reasoning)
- Code generation on a budget? โ Kimi K2.7 Code ($0.95/$4.00, code-optimized)
- General tasks with Chinese language needs? โ Kimi K2.6 ($0.95/$4.00, multilingual)
- Cheapest possible, don't need Chinese optimization? โ DeepSeek V4 Flash ($0.22/$0.66) or Gemini 2.5 Flash-Lite ($0.10/$0.40)
- Best quality at mid-tier price? โ Claude Sonnet 5 ($2.00/$10.00) โ stronger than K3 on most English benchmarks
Moonshot V1 Sunset Warning
If you're still using Moonshot V1 models, they sunset August 31, 2026. Migrate to K2.6 (general) or K2.7 Code (coding) โ both are cheaper and more capable than V1.
The Verdict
The Kimi lineup now covers three price points. K3 is interesting for its 1M context + caching economics โ if your workload has high cache hit rates, effective cost drops dramatically. But at list price, it's more expensive than Claude Sonnet 5 and Gemini 3.1 Pro for most use cases.
The budget models (K2.6, K2.7 Code) are hard to recommend over DeepSeek V4 Flash ($0.22/$0.66) or GPT-5.4 nano ($0.20/$1.25) unless you specifically need Chinese language optimization. They're closer in price to Claude Haiku 4.5, which offers better English quality.
If you're building for Chinese users or need Moonshot's specific strengths, the Kimi models are worth testing. For everything else, the competitive landscape offers cheaper or higher-quality alternatives.
Compare Kimi costs against all 95 models we track
Open the Moonshot Cost Calculator