Speech 02 Turbo
Neural Network
Speech 02 Turbo is a speech synthesis model for text-to-speech with emotional delivery and low latency.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
How it works Speech 02 Turbo?
Speech 02 Turbo is a text-to-audio model: it converts text into spoken speech, conveying emotions and intonations, and supports multiple languages. It is engineered for low latency, making it suitable for applications requiring near-instant voice output: voice assistants, dialogue scenarios, and interactive services. In BotHub, Speech 02 Turbo is accessible without a VPN or foreign card; you pay with a Russian card in rubles only for tokens used, which do not expire. With over 250 neural networks available in one window, you can easily combine voice synthesis with text and visual models. A unified OpenAI-compatible API allows you to integrate speech synthesis into your product and replace the model if needed without rewriting the integration. Requests are encrypted with AES-GCM, and companies have access to contracts, invoices, electronic document management, and an admin panel with limits. Typical tasks include voiceovers for training courses and videos, lines for games and chatbots, audio versions of articles and newsletters, voice notifications in apps, and interface prototypes with live speech.Frequently asked questions about Speech 02 Turbo
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is a speech synthesis model from MiniMax: you send text and receive ready-made audio. It is suitable for voiceovers for videos and reels, audio versions of articles, voice responses in bots and services, training materials, podcasts, and commercials. It works in BotHub from Russia without a VPN or foreign card.
The output is an audio track synthesized from the provided text, which you can immediately download and insert into your edit or feed into your application. The length of a single request is limited to a context of 4095 tokens, so it is more convenient to voice long scripts in parts by semantic blocks.
A separate step-by-step reasoning mode is not indicated in our data, and it is not needed for speech synthesis: the model's task is to turn your text into an audio track, not to analyze it. If you need reasoning on a task, prepare the text in a text model and pass the result here.
Such markings are not present in our data for this model — the generation result is audio, not a text structure. However, you can access it via the unified OpenAI-compatible BotHub API, so integrating voiceover into your service or multi-model pipeline is easy.
Acceptance of files other than text is not indicated in our data: the model expects a text description of what needs to be spoken as input. If your project requires images, video, or speech recognition, switch to a suitable model in the same BotHub interface — you won't have to rewrite the integration.
The speech-02 line from MiniMax is stated to be multilingual, and Russian is supported. It is most convenient to check the quality with your own material: voice a short fragment of the script, evaluate the sound and diction, and then run the full text. In BotHub, you pay for actual usage, without a subscription.