Moonshot: Kimi K3
kimi-k3- Input / Output
- $3.00 / $15.00
- Context
- 1M
- Max output
- 128K
- Modalities
- VisionChat
- Capabilities
- Tool usePrompt CacheReasoning
About Kimi K3
Kimi K3 is Moonshot AI's open-weight flagship, released in July 2026: a 2.8-trillion-parameter mixture-of-experts model with native image and video input and a one-million-token context. Thinking is always on with low, high, or max effort. It uses Kimi Delta Attention and activates 16 of 896 experts per token. Moonshot builds it for long-horizon agents, large codebases, and knowledge work.
Where it works well
- Reasoning effort can be set to low, high, or max, so one model covers quick answers and deep problem solving.
- One-million-token context, aimed at large codebases and long agent histories.
- Native image and video input, so charts, screenshots, and recordings can be part of the task.
- Supports JSON output, tool calling with dynamic loading, prefix continuation, and automatic context caching.
When to choose another model
- Thinking is always on, so trivial turns cost more time than a non-thinking model would.
- A model of this size is heavy to self-host even though the weights are open.
- Text is the only output; use a separate model for images or speech.
Getting started
Create API key
Create a key in Console, then use it with every model on the platform.
Send your first request
Copy the example for your language and run it against the endpoint.
Text to textResponses APIPOST/v1/responsesAPI format:curl https://api.tokenlab.sh/v1/responses \ -H "Content-Type: application/json" \ -H "Authorization: Bearer sk-xxx" \ -d '{ "model": "kimi-k3", "input": "Hello!" }'
Pricing
The Verified price applies to Verified, which costs less on most models. The Official price is the model maker's published price and applies to the more reliable Official route. Auto bills the route that completes the request.
Default
per 1M tokens- Official price
- Input $3.00 / Output $15.00 / Cache read $0.30 / Cache write $3.00 / Text Input $3.00 / Text Output $15.00
- Verified price
- —
- Discount
- —
Cache read
- Official price
- $0.30
- Verified price
- —
- Discount
- —
Cache write
- Official price
- $3.00
- Verified price
- —
- Discount
- —
| Official priceper 1M tokens | Verified priceper 1M tokens | Discount | |
|---|---|---|---|
| Default | Input $3.00 / Output $15.00 / Cache read $0.30 / Cache write $3.00 / Text Input $3.00 / Text Output $15.00 | — | — |
| Prompt cache pricing | |||
| Cache read | $0.30 | — | — |
| Cache write | $3.00 | — | — |
Usage & activity
Success rate is the share of requests that completed. Latency is how long a full response takes; P95 means 95% of requests finished within that time.
Usage & availability
Last 24 hours- Requests
- Success rate
- P95 latency
- Total tokens
- 30-day success rate
- 99.2%
- 30-day requests
- 100+
- Last active
- 11 hours ago
Data is based on aggregate user requests, excluding status checks.
Open in Console
Open Kimi K3 in Console with a prompt ready to edit or send.
Help me try kimi-k3 with a short message at /v1/responses. Show the reply, latency, and cost.
Use cases
Best for- Reasoning
- Vision
Long-horizon agents
Run research or engineering agents that make hundreds of tool calls and need to remember earlier results across a very long session.
Large codebase work
Navigate and change a big repository by loading much of it into context rather than relying only on retrieval.
Video and chart analysis
Ask questions about a recorded demo or a dense chart and receive findings that cite what is on screen.
Knowledge-work automation
Draft reports and analyses from large document sets, using tools to fetch data and check numbers.
Prompt examples
Review the attached 600-page filing and list every related-party transaction with its page number.
Read this repository and explain how a request flows from the router to the database, naming each file.
Watch this product demo video and write the bug report for the step where the form resets.
This model has conditional pricing. Monthly token totals alone cannot produce a reliable estimate; use the detailed pricing for the request specification, cache, and applicable time window.
FAQ
What is Kimi K3?
Moonshot AI's open-weight flagship model, with 2.8 trillion parameters in a mixture-of-experts design. It reads text, images, and video, has a one-million-token context, and thinks before answering with adjustable effort.
What are the reasoning effort levels?
Low, high, and max. Thinking is always on, and effort sets how much the model deliberates. Use low for routine requests and max for hard debugging, planning, or analysis.
Is Kimi K3 open source?
Moonshot published the weights after the API launch, so you can self-host if you have the hardware. Calling it through an API avoids that cost and lets you switch models without redeploying.
Does it support video input?
Yes. Moonshot describes native support for images and video, which suits screen recordings, demos, and charts alongside text instructions. The output stays text, so findings come back as written answers.
How is it different from Kimi K2.6?
K3 is the newer and much larger model, with a one-million-token window, configurable effort levels, and more efficient scaling by Moonshot's account. K2.6 has a shorter context and no effort setting.
How much does Kimi K3 cost?
On TokenLab, Kimi K3 costs Input $3.00 / Output $15.00 per 1M tokens. The pricing table above shows the full breakdown. Rates depend on the billing unit, specification, and usage. Compare matching conditions in the model's detailed pricing; a single rate does not determine the total cost.
What are the context window and output limit of Kimi K3?
Kimi K3 accepts up to 1,048,576 tokens of context and returns up to 131,072 tokens in one response.
Which endpoint should Kimi K3 use?
Use https://api.tokenlab.sh/v1/responses for Kimi K3. The request example below shows the matching code shape.
Which operations does Kimi K3 support?
Kimi K3 supports Text to text. Select an operation above to see its endpoint and request example.
Compare Kimi K3
Sources
Reviewed Oct 2, 2026