Kimi K2
Neural Network
Kimi K2 API: a large Moonshot AI language model using Mixture-of-Experts architecture: 1 trillion parameters, 32 billion active.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Kimi K2?
Kimi K2 Instruct is a large language model by Moonshot AI built on the Mixture-of-Experts architecture: it has a trillion parameters, but only about 32 billion are activated per pass, meaning the request processes only a part of the network. This is an instruction-tuned version designed for tasks in chat or via API. In BotHub, the model is accessible from Russia without a VPN or foreign card: you pay in rubles for tokens used, with no subscription, and unused tokens do not expire. Other aggregator models work alongside in the same window, so you can compare answers and switch models without rewriting your integration—the unified API is the same for all models, correspondence is encrypted via AES-GCM, and teams have access to contracts, invoices, and an admin panel with limits. It is useful for analyzing and refactoring code, writing and proofreading documentation, consolidating disparate data into structured reports, preparing support response drafts, or building a multi-model service where Kimi K2 handles part of the requests.Frequently asked questions about Kimi K2
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
Kimi K2 is a Moonshot AI text model with a 131,072 token context. It accepts text, documents, and images as input and can call functions. Use it for working with code, analyzing long files and reports, analytics, and agent scenarios where you need to keep a lot of context in memory.
The single response limit is 98,304 tokens, which is dozens of pages of text. This capacity is enough for a large code module, a detailed report, or an extensive document analysis without cutting off in the middle. If the response ends early, just ask it to continue from where it stopped.
There is no specific note about hidden reasoning in our data, but that doesn't mean the model can't do it. It is easiest to control this via the prompt: ask it to break down the task step-by-step, list options, and only then provide the final conclusion.
Yes, function calling and structured JSON output are confirmed. The model can be connected to external tools, search, or databases to receive responses in a strict schema. In BotHub, this works via a unified OpenAI-compatible API, so switching models won't require rewriting your integration.
Images and documents, yes, you can send them as input and ask it to analyze the content. There is no note about audio and video in our data, so for such tasks, it is more convenient to use specialized models: in BotHub, they open in the same window without switching services.
There is no note about languages in our data, so we won't make promises on behalf of the model. You can check this in a minute: give Kimi K2 your real text and compare it with the response of another model in the same BotHub window. Payment is based on actual token usage, without a subscription.