Mistral Nemo

Neural Network

Mistral Nemo is a 12B parameter multilingual language model from Mistral and NVIDIA with a 128k token context window.

Main

/

Models

/

Mistral Nemo
16 384

Max answer length

(in tokens)

128 000

Context size

(in tokens)

2,24 ₽

Prompt cost

(per 1M tokens)

3,54 ₽

Answer cost

(per 1M tokens)

*Prices are shown for API usage via ECO providers.
bothub
BotHub: Try neural networks for freebot

Caps remaining: 0 CAPS
Code example and API for Mistral NemoWe offer full access to the OpenAI API through our service. All our endpoints fully comply with OpenAI endpoints and can be used both with plugins and when developing your own software through the SDK.Create API key
Javascript
Python
Curl
illustaration

How it works Mistral Nemo?

Mistral Nemo is a 12 billion parameter language model developed by Mistral in collaboration with NVIDIA. With a 128k token context window, the model handles long dialogues and large documents in their entirety, while its multilingual capabilities allow it to work with texts in various languages. Its compact size makes it a convenient option where response speed is critical for high-volume requests. In BotHub, the model is accessible without a VPN or foreign card: you pay with a local card in rubles based on actual token usage, not a subscription. Over 250 other neural networks are available in the same window—compare answers or switch with a few clicks. A unified OpenAI-compatible API eliminates the need to rewrite integrations; chats are encrypted via AES-GCM, and history is not saved. For teams, we offer contracts, invoicing, and an admin panel with limits. Use cases include: analyzing long reports and correspondence, summarizing and structuring notes, drafting emails and descriptions, multilingual dialogues in chatbots and support, and translating meaning across languages for distributed teams.

Frequently asked questions about Mistral Nemo

Can I use Mistral Nemo results for commercial purposes?

You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.

What can mistral-nemo do and what tasks is it suitable for?

mistral-nemo works with text: it answers questions, edits and summarizes documents, writes and analyzes code, and helps with analytics. It accepts images and documents as input and supports function calling, making it suitable for both conversational scenarios and integration into services via API.

How much text can mistral-nemo write in a single response?

Up to 16,384 tokens per response—this is the volume of a large article or several pages of code. The context window is 128,000 tokens, so the model will hold a long document and the entire correspondence, and if necessary, you can continue the response with the next request.

Can mistral-nemo reason before answering?

A separate reasoning mode for mistral-nemo is not specified in our data, so we won't promise it. In practice, a simple trick helps: ask the model in the prompt to break down the task step-by-step—the model will unfold the logic directly in the response, and you will see the solution process.

Does mistral-nemo support function calling and JSON output?

Yes. Function calling and structured JSON output are supported: you describe the tools, the model selects the necessary one and returns arguments according to the specified schema. This is convenient for agents, document parsing, and integrations. In BotHub, all this is available via a unified OpenAI-compatible API.

Can I upload images, audio, or video to mistral-nemo?

Images and documents—yes, they can be sent as input along with text; the model will analyze the content and respond with text. Audio and video are not supported in our data, and image or sound generation is not provided: for this, switch to a specialized model in the same window.

How well does mistral-nemo understand Russian?

The language of prompts and responses is not specified in our data, so we won't make any claims—it's easier to test it with your own task, which takes a couple of minutes. If the result is not satisfactory, switch to another model in the same window without rewriting the integration, and pay only for the tokens used.

Support ServiceOpen from 10:00 to 18:00 MSK