How it works TBank VoiceKit STT?
Frequently asked questions about TBank VoiceKit STT
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
voicekit-stt is a speech recognition model from T-Bank VoiceKit: it converts audio to text, including long recordings. It is suitable for transcribing meetings, interviews, calls, and podcasts, as well as preparing notes and subtitle drafts. Accessible from Russia without a VPN or foreign card, payment in rubles.
In a single response, the model provides up to 4096 tokens of text, with a context of 4095 tokens. This is enough to transcribe a fragment of a recording; if the audio is long, split it into parts and merge the results or process them in batches via the API.
Step-by-step reasoning is not indicated in our data, and it is not the main focus for this task: voicekit-stt is designed for accurate speech transcription, not for reasoning. If you need analysis, conclusions, or a summary of the finished text, pass the transcription to any text model in BotHub.
Function calling and structured JSON output are not indicated in our data. The output is recognized text. If you need structure — fields, tags, tables — send the transcription to a text model via the unified OpenAI-compatible BotHub API without rewriting the integration.
Audio — yes, this is the main input: the model accepts recordings, including large files. Image and video input is not indicated in our data, so count on audio tracks. If you need to work with images or video, switch to another model in the same window.
This is a development by the Russian vendor T-Bank VoiceKit, and the service was created for Russian speech recognition tasks. We do not publish exact quality metrics, so the best way is to run your own recordings: meetings, interviews, calls — and compare the transcription with other models in BotHub.