[Moonshot AI](/ai-gateway/models/labs/moonshotai)

# Kimi K2 Thinking

Kimi K2 Thinking adds extended chain-of-thought (CoT) reasoning to the K2 architecture, supporting many sequential tool calls for agentic workflows through AI Gateway. Your use is subject to Moonshot AI's [Terms](https://platform.moonshot.ai/docs/agreement/modeluse.en-US) & [Privacy](https://platform.moonshot.ai/docs/agreement/userprivacy.en-US) Policies.

ReasoningTool UseImplicit Caching

[Use with AI Gateway](https://vercel.com/d?to=%2F%5Bteam%5D%2F%7E%2Fai%3Futm_source%3Dgateway-model-page%26utm_campaign%3Dai-gateway-models&title=Get+Started+with+Vercel+AI+Gateway) [View docs](https://vercel.com/docs/ai-gateway)

AI SDKChat CompletionsMessagesResponses

```
1import { streamText } from 'ai'
2

3const result = streamText({
4  model: 'moonshotai/kimi-k2-thinking',
5  prompt: 'Why is the sky blue?'
6})
```

[Read docs](https://vercel.com/docs/ai-gateway/sdks-and-apis/ai-sdk)

[Overview](/ai-gateway/models/kimi-k2-thinking) [API](/ai-gateway/models/kimi-k2-thinking/api) [About](/ai-gateway/models/kimi-k2-thinking/about) [Providers](/ai-gateway/models/kimi-k2-thinking/providers) [Throughput](/ai-gateway/models/kimi-k2-thinking/throughput) [Latency](/ai-gateway/models/kimi-k2-thinking/latency) [Uptime](/ai-gateway/models/kimi-k2-thinking/uptime) [Status](/ai-gateway/models/kimi-k2-thinking/status) [Similar](/ai-gateway/models/kimi-k2-thinking/similar) [FAQ](/ai-gateway/models/kimi-k2-thinking/faq)

## [Copy link to heading](#playground)Playground

Try out Kimi K2 Thinking by Moonshot AI. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.

![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75)Kimi K2 Thinking

![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=96&q=75)

Kimi K2 Thinking

## [Copy link to heading](#providers)Providers

Route requests across multiple providers. Copy a provider slug to set your preference. Visit the [docs](/docs/ai-gateway/provider-options) for more info. Using a provider means you agree to their terms, listed under Legal.

| Provider |
| --- |

| Context | Max Output | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) [DeepInfra](/ai-gateway/models/providers/deepinfra) Legal:[Terms](https://deepinfra.com/terms)•[Privacy](https://deepinfra.com/privacy) | 216K | 216K | 0.8s | 56tps | $0.47/M | $2/M | Read:$0.14/M Write:— | — |  |  |  | 11/06/2025 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

## [Copy link to heading](#throughput)Throughput24 hours

1W

1D

P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#latency)Latency24 hours

1W

1D

P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/metrics) for more info.

## [Copy link to heading](#uptime)Uptime24 hours

1W

1D

1H

Direct request success rate on AI Gateway and per-provider. Visit the [docs](https://vercel.com/docs/ai-gateway/models-and-providers/uptime) for more info.

1W

1D

1H

## [Copy link to heading](#more-models-by-moonshot-ai)More models by Moonshot AI

All

Text

Code

Video Input

| Model |
| --- |

| Context | Latency | Throughput | Input | Output | Cache | Web Search | Capabilities | Providers | ZDR | No Training | Release Date |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |

| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi\-k3-fast](/ai-gateway/models/kimi-k3-fast) | 1M | 1.0s | 107tps | $4.50/M | $22.50/M | Read:$0.45/M Write:— | — | +2 | ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![morph logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmorph.png&w=48&q=75) |  |  | 07/27/2026 |  |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k3](/ai-gateway/models/kimi-k3) | 1M | 0.5s | 221tps | $2.90/MFast $4.50/M | $14/MFast $22.50/M | Read:$0.29/M Write:— | — | +2 | ![alibaba logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Falibaba%2520cloud.png&w=48&q=75) ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![digitalocean logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdigitalocean.png%3Fv%3D1784246703962&w=48&q=75) +7 |  |  | 07/16/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.7-code-highspeed](/ai-gateway/models/kimi-k2.7-code-highspeed) | 262K | 0.6s | 96tps | $1.90/M | $8/M | Read:$0.38/M Write:— | — | +2 | ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) |  |  | 06/15/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.7-code](/ai-gateway/models/kimi-k2.7-code) | 262K | 0.5s | 185tps | $0.74/MFast $1.90/M | $3.50/MFast $8/M | Read:$0.15/M Write:— | — | +2 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![deepinfra logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fdeepinfra.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) +1 |  |  | 06/12/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.6](/ai-gateway/models/kimi-k2.6) | 262K | 0.3s | 128tps | $0.95/M | $4/M | Read:$0.16/M Write:— | — | +1 | ![baseten logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fbaseten.png&w=48&q=75) ![fireworks logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Ffireworks.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) +1 |  |  | 04/20/2026 |  |
| ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) [moonshotai/kimi-k2.5](/ai-gateway/models/kimi-k2.5) | 262K | 0.5s | 118tps | $0.60/M | $3/M | Read:$0.10/M Write:— | — | +1 | ![bedrock logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Famazon%2520bedrock.png&w=48&q=75) ![moonshotai logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fmoonshotai.png&w=48&q=75) ![novita logo](/vc-ap-vercel-marketing/_next/image?url=https%3A%2F%2F7nyt0uhk7sse4zvn.public.blob.vercel-storage.com%2Fdocs-assets%2Fstatic%2Fdocs%2Fai-gateway%2Flogos%2Fnovita.png&w=48&q=75) |  |  | 01/26/2026 |  |

## [Copy link to heading](#about-kimi-k2-thinking)About Kimi K2 Thinking

Standard language models produce answers directly. Input goes in, output comes out, and whatever reasoning occurred stays invisible. Kimi K2 Thinking changes the output structure. Before generating its final answer, the model produces an explicit chain-of-thought (CoT) trace: a written record of how it decomposes the problem, what options it considers, and how it reaches its conclusion.

This isn't a prompting trick. The thinking behavior is trained into the model. When K2 Thinking encounters a hard problem, its reasoning trace can run for hundreds or thousands of tokens as the model works through sub-problems, backtracks from dead ends, and synthesizes intermediate results. The final answer follows the trace.

Two practical consequences follow. First, step-by-step decomposition helps on problems that benefit from it: multi-step mathematical proofs, algorithmic design, and debugging sessions where the root cause isn't obvious. Second, the reasoning trace is also an output you can log, audit, or use in evaluations.

K2 Thinking supports long chains of sequential tool calls within a single agentic session. The model reasons about what tool to call next, observes the result, reasons about the implications, and continues. It maintains coherent task state across more interaction steps than many non-thinking models handle.

The model is open source under Moonshot AI's license terms.

Kimi K2 Thinking is available through AI Gateway at $0.47 per million input tokens and $2 per million output tokens.

## [Copy link to heading](#what-to-consider-when-choosing-a-provider)What To Consider When Choosing a Provider

- Configuration: Reasoning traces increase output length, so budget planning should account for higher output token use relative to non-thinking K2 variants. Completions support up to 216.1K tokens per request.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the [documentation](https://vercel.com/docs/ai-gateway/security-and-compliance/zdr) for details.
- Authentication: AI Gateway authenticates requests using an [API key](https://vercel.com/docs/ai-gateway/authentication-and-byok#api-key-authentication) or [OIDC token](https://vercel.com/docs/ai-gateway/authentication-and-byok#oidc-token-authentication). You do not need to manage provider credentials directly.

## [Copy link to heading](#when-to-use-kimi-k2-thinking)When to Use Kimi K2 Thinking

### Best for

- Visible model reasoning: Problems where seeing the model's work matters — debugging complex logic, validating mathematical derivations, auditing decisions
- Algorithmic exploration: Multi-step design where the model must explore and eliminate approaches before settling on a solution
- Long tool-call chains: Agentic sessions requiring sequential tool calls with coherence across the full chain
- Evaluation and red-teaming: Workflows where reasoning traces surface failure modes and edge cases

### Consider alternatives when

- Straightforward tasks: Standard Kimi K2 is faster and cheaper for direct-answer tasks that don't benefit from deliberation
- Hard latency constraints: Reasoning traces add significant generation time
- Output cost sensitivity: Thinking traces can multiply output length by 3 to 10x
- Speed-optimized reasoning: Kimi K2 Thinking Turbo trades some reasoning depth for lower latency

## [Copy link to heading](#conclusion)Conclusion

Kimi K2 Thinking restructures model output around explicit reasoning. For problems that reward deliberation, the visible chain-of-thought adds an auditable record of the model's logic. Long chains of sequential tool calls extend this into agentic workflows. Reserve it for tasks where the thinking trace earns its token cost. Use non-thinking variants for everything else.