Kimi K2 Thinking
Neural Network
Kimi K2 Thinking is an open reasoning model by Moonshot AI for agentic and multi-step tasks, available via Kimi K2 Thinking API.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Kimi K2 Thinking?
Kimi K2 Thinking is an open reasoning model by Moonshot AI, continuing the K2 series on a trillion-parameter Mixture-of-Experts architecture. It focuses on agentic scenarios and long reasoning chains, designed for tasks requiring multi-step sequences rather than single answers. In BotHub, the model is accessible without a VPN or foreign card; payment is in rubles via Russian cards, based on usage rather than subscription. With over 250 models in the same window, you can compare Kimi K2's responses with others and switch instantly. For integrations, we offer a unified OpenAI-compatible API: one connection, switch models via request parameter. Analysts can use it for multi-step insights on large datasets, researchers for hypothesis development, developers as an agent reasoning core, and product teams for prototypes with complex logic.Frequently asked questions about Kimi K2 Thinking
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
Kimi K2 Thinking is a reasoning model from moonshotai. It works with text and documents, accepts images, and calls external functions. A 262,144-token context allows keeping a large repository, technical specifications, or a collection of files in the dialogue, making the model convenient for code, analytics, and agentic scenarios.
The output limit is 235,929 tokens per response. This is enough for a voluminous technical document, a large code module, or a detailed analysis with calculations. In practice, the length depends on the task and your prompt: if you need the maximum, explicitly ask for a detailed and complete answer.
Yes, this is a key feature of the thinking version: before answering, the model builds a chain of reasoning and only then provides the result. This mode significantly helps with math, logic, code debugging, and multi-step tasks where a correct answer is more important than a fast one.
Yes. The model supports function calling and structured JSON output, so it can be connected to your tools, databases, and external services. In BotHub, this is available via a unified OpenAI-compatible API — you can switch models without rewriting the integration.
Images and documents — yes, they can be sent as input along with text: analyze a screenshot, diagram, table, or file. Audio and video support is not noted in our data. Image or sound generation does not apply to this model — the output is text only.
There are no separate language benchmarks in our data, so we won't promise anything specific — it's easier to test the model on your task, especially since you can switch to another one in BotHub with one click. Access from Russia is without VPN, payment in rubles.