Voxtral Small 24B 2507
Neural Network
Voxtral Small 24B is a Mistral model with audio input for transcription, speech translation, and understanding audio recordings.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Voxtral Small 24B 2507?
Voxtral Small 24B is a Mistral development: it is an evolution of Mistral Small 3 with added audio input processing, while retaining the text capabilities of the base model. The model accepts audio and handles speech transcription, translation, and understanding of recording content, while remaining a standard text model for familiar tasks. In BotHub, you work with it without a VPN or foreign card: payment with Russian cards in rubles and pay-as-you-go, tokens do not expire. With over 250 models in one window, you can compare results with another model in a few clicks, and a unified OpenAI-compatible API allows you to connect Voxtral Small to your service and switch later without rewriting the integration; requests are encrypted with AES-GCM. It is useful if you need to turn a call or interview recording into text, summarize a podcast or lecture, process customer audio requests in support, translate speech from a recording, or embed voice input into your own application.Frequently asked questions about Voxtral Small 24B 2507
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
Voxtral Small 24B from Mistral AI works with text and documents: it answers questions, summarizes and structures file content, and helps with correspondence, code, and analytics. It supports function calling, making it suitable for integration into services via the unified OpenAI-compatible BotHub API.
Up to 26,214 tokens in a single response — enough for a lengthy article, detailed document analysis, or a large code snippet. The total context window is 32,000 tokens, which includes your request with files and the response itself.
A separate step-by-step reasoning mode is not noted in our data for this model. However, you can explicitly ask in the prompt to break down the task step-by-step, list the conditions, and verify the output — in practice, this significantly improves the accuracy of the analysis.
Yes. The model supports function calling and structured JSON output: you describe the tools, it selects the necessary one and returns arguments in the specified schema. Convenient for agents, parsing documents into fields, and integrations — requests go through the OpenAI-compatible API.
According to our data, text and documents are accepted as input — you can upload a file in its entirety, and the model will analyze its content. Input of images, audio, and video is not separately noted, so for such tasks, it is more convenient to select a specialized model in BotHub in the same window.
We do not have data on prompt and response languages, so we will not make any promises here — please test it with your own scenario. You can compare several models on the same request directly in BotHub, without a VPN or foreign card, with payment in rubles.