Minimax Speech 2.8 Turbo
Neural Network
Minimax Speech 2.8 Turbo is a speech synthesis model from MiniMax with voice cloning and emotion control for text-to-speech.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
How it works Minimax Speech 2.8 Turbo?
Minimax Speech 2.8 Turbo is a MiniMax solution for text-to-speech: the model reads text with a natural voice, conveying intonation and emotion, while voice cloning allows you to establish a recognizable sound for your project. With support for over 40 languages, one model covers multilingual materials. In BotHub, it is accessible without a VPN or foreign card: pay with a Russian card in rubles only for actual usage, and unused balance does not expire. With over 250 models in the same window, you can easily build a complete workflow: prepare text in one, synthesize audio here. For production tasks, there is a unified OpenAI-compatible API. Companies have access to contracts, invoices, electronic document management, and an admin panel with limits, with traffic encrypted via AES-GCM. Useful for narrating educational courses and podcasts, creating voices for games and assistants, preparing audio versions of articles and newsletters, localizing videos for different markets, or testing voice interfaces before studio recording.Frequently asked questions about Minimax Speech 2.8 Turbo
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
speech-2.8-turbo is a speech synthesis model from Minimax: you provide text, and receive audio output. Suitable for videos, podcasts, audio versions of articles, IVR voice responses, and educational materials. Text is submitted within a 4095-token context.
You receive a finished audio track as output, not text. The specific file format and parameters are set in the API request; they are not fixed in our data. The output limit is 4096 tokens per request, so it is more convenient to narrate long scripts in parts.
There is no indication of reasoning before answering in our data, and it is not required for speech synthesis: the model works with finished text. If you need to think through a script or lines first, prepare them with a text model in the same BotHub window, then send them to speech-2.8-turbo.
Function calling and structured JSON output are not noted for speech-2.8-turbo in our data — it is a speech synthesis model, not an agent tool. You can access it via the unified OpenAI-compatible BotHub API, and keep call logic on your application side.
File acceptance beyond text is not noted in our data: speech-2.8-turbo works based on text descriptions and outputs audio. If you need to recognize a recording, analyze an image or video, select a suitable model in the same window — integration via the common API remains the same.
The list of supported languages is not specified in our data, so we will not promise accuracy in Russian — test it with your own text, it's quick. BotHub itself is fully in Russian: interface, support, and payment with Russian cards without VPN or foreign cards.