Qwen 2.5 72B Instruct
Neural Network
Qwen 2.5 72B Instruct is a Qwen language model with expanded knowledge and improved coding capabilities, available via chat and API.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Qwen 2.5 72B Instruct?
Qwen 2.5 72B Instruct is the instruct version of the 72-billion parameter Qwen2.5 language model. Compared to the previous Qwen2 generation, this model has significantly increased its knowledge base and coding capabilities, making it ideal for tasks requiring precise analysis and clear technical answers. In BotHub, it is accessible without a VPN or foreign card: you pay with a Russian card in rubles based on actual token usage, not a subscription, and the balance does not expire. Over 250 models are available in the same window, allowing you to easily switch between them and compare answers for the same prompt. A unified OpenAI-compatible API lets you switch models in your project without rewriting the integration. Dialogues are transmitted with AES-GCM encryption and are not saved. It is useful for writing and refactoring code, understanding third-party projects, debugging, preparing technical documentation, gathering information on complex topics, and automating request processing and correspondence via API. BotHub provides corporate access for teams: contracts, invoicing, EDI, admin panel, and employee limits.Frequently asked questions about Qwen 2.5 72B Instruct
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is an instruct language model from Qwen with 72 billion parameters. It is suitable for writing and editing text, working with code (generation, refactoring, debugging), analyzing documents, answering questions about uploaded files and images, as well as for STEM tasks and logical reasoning.
The single response limit is 16,384 tokens, which is approximately 25–30 pages of text. The context window is 131,072 tokens, so you can upload a voluminous document or a large fragment of a codebase into the dialogue and receive a detailed analysis in its entirety.
A separate hidden reasoning mode is not noted in our data. However, you can ask the model to write out the solution step-by-step directly in the prompt — for math, logic, and code analysis, this approach works well and provides a transparent train of thought.
Yes, the model supports function calling and structured JSON output. This is convenient for agents, integrations with external services, and data parsing. A unified OpenAI-compatible API is available in BotHub, so you can connect it without rewriting existing integrations.
Text, images, and documents are accepted as input: you can ask to describe a picture, extract data from a screenshot, or analyze a long file. This model does not process audio or video — BotHub has separate models for those, and you can switch between them in one window.
The Qwen 2.5 series was trained as multilingual, and Russian is among the supported languages, so the model confidently handles dialogue, translation, and processing of Russian-language documents. You can easily check the quality on specific tasks yourself by comparing answers with other models in BotHub.