Qwen2.5 VL 72B Instruct
Neural Network
Qwen2.5 VL 72B Instruct is a multimodal Qwen model for image analysis: objects, text, charts, icons, and layout markup.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Qwen2.5 VL 72B Instruct?
Qwen2.5 VL 72B Instruct is a vision-language model from the Qwen family that accepts both text prompts and images. The model confidently recognizes common objects like flowers, birds, fish, and insects, and parses content within images: text fragments, charts, icons, diagrams, and layout structures. In BotHub, access is available without VPN or foreign cards: pay with Russian cards in rubles only for tokens actually used, which do not expire. Access over 250 other neural networks in one window, easily switching between them to compare responses. A unified OpenAI-compatible API allows you to connect the model to your service and switch later without rewriting integrations. Data is encrypted via AES-GCM and not stored. Companies have access to contracts, invoices, electronic document management, and an admin panel with limits. Simple scenarios: extract numbers from charts for reports, pull text from screenshots or scans, describe photo content for catalogs, analyze interfaces or layouts by elements, or recognize objects in field photos.Frequently asked questions about Qwen2.5 VL 72B Instruct
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is a Qwen vision-language model: it reads text, parses images and documents, and answers questions about their content. Suitable for analyzing screenshots, diagrams, tables, and scans, extracting data into structured formats, working with code, and connecting tools via function calling. Context: 32,000 tokens.
Up to 115,200 tokens per response according to our data, which is the volume of a large article or entire documentation. The context window is 32,000 tokens, so count the prompt and attached files together, and break tasks into parts for long materials.
A separate reasoning mode is not noted in our data, but that doesn't mean it doesn't exist. In practice, it helps to ask the model to break down the task step-by-step: this makes answers for logic, calculations, and document analysis more accurate. Test it with your scenario.
Yes. Function calling and structured JSON output are supported: the model can access your tools and return answers in a specified schema. This is useful for extracting fields from scans and invoices, integrations, and agent scenarios. In BotHub, everything works via a unified OpenAI-compatible API.
Images and documents, yes, they can be sent as input along with text: the model will describe the image, find the required fragment in a screenshot, or parse a table or contract. Audio and video input is not noted in our data. BotHub has separate models for such tasks.
We do not have data on prompt and output languages, so we cannot promise specific quality. Qwen2.5-VL belongs to the multilingual Qwen line, so the model will understand Russian queries. It is best to test it with your task and compare it with another model—you can switch in one window in BotHub.