How it works Whisper 1?
Frequently asked questions about Whisper 1
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
OpenAI's whisper-1 converts audio to text. You upload a recording and get a transcript. It is suitable for interviews, podcasts, calls, lectures, voice messages, and subtitle drafts. It is convenient when you need to quickly turn hours of conversation into readable text for searching and editing.
Up to 4096 tokens per response — this is approximately several pages of transcription; the model's context is 4095 tokens. It is better to cut long recordings into fragments and send them one by one, then stitch the result together: this way, nothing will be cut off in the middle of a sentence.
In our data, a separate reasoning mode for whisper-1 is not noted, and it is not needed for this task: the model recognizes speech, it does not build chains of thought. If you need an analysis of the transcript, pass the text to a text model in the same BotHub window.
In our data, function calling and strict JSON output are not noted for whisper-1. Access is via the unified OpenAI-compatible BotHub API: you send a file and receive a transcript, and you can structure it on your side or in the next step of the pipeline.
Audio — yes, this is the model's main input: whisper-1 turns the audio track into text. Images are not noted as input in our data. Video is handled simply: extract the audio track with any converter and send it for recognition.
whisper-1 is a multilingual speech recognition model; it understands and transcribes Russian speech. Quality depends more on the recording itself: clear sound, a single speaker, and a calm pace produce noticeably more accurate text than a noisy room or a strong accent. Check it with your own file.