Speech 02 HD
Neural Network
Speech 02 HD by MiniMax: high-quality speech synthesis for text-to-speech, audiobooks, and video voiceovers.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
How it works Speech 02 HD?
Speech 02 HD is a speech synthesis model by MiniMax: input text, output a ready-made audio track with emotional delivery. The model is designed for high sound quality and multi-language support, making it ideal for tasks requiring natural intonation rather than mechanical reading: voiceover tracks, audiobooks, and long narrative formats. Via BotHub, Speech 02 HD is accessible from Russia without a VPN or foreign card; payment is made with Russian cards in rubles, based on actual usage, with no subscription. Over 250 other neural networks are available in the same window, and a unified OpenAI-compatible API allows you to integrate voiceover into your service and switch models if needed without rewriting the integration, with data transmitted in encrypted form. Common use cases: narrating audiobooks and podcasts, creating voice tracks for videos, ads, and presentations, recording lessons and online courses, generating lines for games and bots, and preparing voice responses and announcements for customer services.Frequently asked questions about Speech 02 HD
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
speech-02-hd is a speech synthesis model by MiniMax: it turns text descriptions into ready-made audio. Suitable for voiceovers for videos and podcasts, audio versions of articles, voice prompts in apps, and educational materials. The model accepts up to 4095 text tokens per request.
Our data confirms one thing: you get audio as output — speech synthesized from your text. The 'hd' index in the name refers to MiniMax's high-quality branch. Specific file parameters are set on the API side, so it's easier to check the result once with your own snippet.
Separate reasoning before answering is not noted in our data, and it is usually not needed for voiceover: the model does not solve tasks but reads out ready-made text. Assemble the script and logic using a text model in BotHub, and connect speech-02-hd at the final step.
Function calling and structured JSON output are not noted in our data: the model's task is to provide audio, not a text response in a schema. If you need tools and strict formats, place a text model in front: BotHub's unified OpenAI-compatible API links them without rewriting the integration.
Accepting files other than text is not noted in our data, so expect a text-to-audio scenario: you send text, you get voiceover. If you need to analyze an image, recording, or video, models with other capabilities are available nearby in BotHub, and switching happens in one window.
Our data does not specify support for specific languages, so we won't make promises for the model: run a short snippet and listen to how the result sounds. Access to speech-02-hd from Russia works without a VPN or foreign card, payment is via Russian cards, charged based on actual usage.