MiniMax Speech 2.6 Turbo
Neural Network
MiniMax Speech 2.6 Turbo is a speech synthesis model from MiniMax: 300+ voices, emotional intonation, and low latency.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
How it works MiniMax Speech 2.6 Turbo?
MiniMax Speech 2.6 Turbo is a low-latency speech synthesis model from MiniMax with emotional delivery. It features over 300 voices and multilingual support, allowing you to voice the same phrase with different timbres and match the intonation to the material's mood—from calm narration to expressive advertising delivery. Via BotHub, the model is accessible without a VPN or foreign card, with pay-as-you-go billing in rubles. All models—voice, text, and image—are gathered in one window, making it easy to integrate voiceovers into a pipeline with script or video generation. Technical teams can connect MiniMax Speech to services via a unified OpenAI-compatible API. For companies, we offer contracts, invoices, electronic document management, an admin panel, and usage limits. Typical tasks include voiceovers for videos, ad creatives, and training courses, audio versions of articles and landing pages, voice greetings and IVR, audiobooks and podcasts, as well as voice assistant prototypes where fast audio response is crucial.Frequently asked questions about MiniMax Speech 2.6 Turbo
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
speech-2.6-turbo is a speech synthesis model from MiniMax: it converts text into spoken speech. It is suitable for voiceovers for videos and clips, audiobooks, podcasts, voice responses in bots, and telephony. Available in BotHub along with other models in a single interface, without a VPN or foreign card.
You receive finished audio as output: specific formats and sample rates are set by the API request; they are not fixed in our data. Keep in mind the input text length limitation—the model's context is 4095 tokens, so it is better to break long scripts into fragments and stitch them together.
This is a speech synthesis model, not a text assistant, so step-by-step reasoning before answering is not required. There is no specific note about reasoning in our data—we cannot confirm or deny this capability and will not speculate. For details, please refer to the MiniMax documentation.
Our data does not indicate support for function calling or structured JSON output, so we will not claim that speech-2.6-turbo supports this. The model is marketed as a text-to-audio generator, and its primary task is voiceover. See the current documentation for the exact list of parameters.
According to our data, input is limited to text: there is no indication of image, audio, or video input, so we do not claim that such input exists or does not exist. If you need to work with a reference voice, check the parameter list in the BotHub interface.
MiniMax speech models are designed for multilingual synthesis, and Russian is among the supported languages, although there is no separate quality assessment for Russian speech in our data. You should check the pronunciation of names, abbreviations, and numbers with your own text. Access via BotHub is available from Russia, with payment via Russian cards.