How it works TTS 1?
Frequently asked questions about TTS 1
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is a speech synthesis model from OpenAI: you send text and receive a ready-made audio track. It is suitable for video and clip voiceovers, audio versions of articles, voice responses in bots and services, and voicing notifications and educational materials.
The output is an audio file with synthesized speech that can be immediately inserted into a video, newsletter, or application. In a single request, the model processes text within a context of about 4095 tokens; it is convenient to voice long materials in parts.
A separate reasoning mode is not noted in our data, and it is not needed for speech synthesis: the model's task is to voice the submitted text. If you need logic, reasoning, or STEM, prepare the text in a text model and pass the result here.
Function calling and structured output are not noted for this model in our data: it returns audio, not a text response according to a schema. It is convenient to build logic with function calling on a text model and connect tts-1 as a voiceover step via the unified BotHub API.
Accepting files other than text is not noted in our data — work is based on the text description you submit for voiceover. If your scenario requires images or video, connect another model in the same BotHub interface and build a chain.
We do not have data on languages, so it is more honest to check in practice: send a short fragment and listen to the result. Payment in BotHub is based on actual token usage, without a subscription, so the test will cost minimal expenses, and unused tokens do not expire.