Qwen3 VL 235B A22B Instruct
Neural Network
Qwen3 VL 235B A22B Instruct is a Qwen multimodal model for analyzing images, videos, documents, charts, and tables.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Qwen3 VL 235B A22B Instruct?
Qwen3 VL 235B A22B Instruct is an open-weights multimodal model from the Qwen team that combines text generation with image and video understanding. The Instruct version is optimized for vision and language tasks: answering questions about images, analyzing documents, and reading charts and tables. You can submit both text prompts and visual materials to receive a coherent analysis. In BotHub, the model is available without VPN or foreign cards, with payment in rubles via Russian cards for actual token usage, which do not expire. Over 250 models are available in one window for easy comparison, and a unified OpenAI-compatible API allows you to switch models without rewriting integrations. Chats are encrypted via AES-GCM, data is not stored, and corporate options include contracts, invoices, EDI, and an admin panel with limits. Use cases include extracting data from scans and PDFs, describing photo or frame content, converting charts to tables, analyzing support ticket screenshots, and finding specific fragments in video archives.Frequently asked questions about Qwen3 VL 235B A22B Instruct
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is a Qwen model that works with text, images, and documents. It is suitable for analyzing screenshots, diagrams, and scans, parsing long materials, writing and refactoring code, step-by-step reasoning, and connecting external tools via function calling.
The single response limit is 32,768 tokens, enough for a lengthy article, detailed analysis, or a large code module. The total context window is 131,072 tokens, which includes your request with files and the response itself. It is more convenient to receive very large texts in parts.
Yes, the model has a reasoning mode: it breaks down tasks step-by-step before answering. This significantly helps with math, logic, data analysis, and code debugging. Ask it to explain the solution step-by-step for a more precise answer, though it may take slightly longer.
Yes, both function calling and structured JSON output are supported. You can connect the model to your services, databases, and search via the unified BotHub OpenAI-compatible API, and receive responses strictly according to a specified schema—useful when the result goes directly into code.
It accepts images and documents: screenshots, photos, diagrams, tables, and scans. The model will describe the content, extract data, and answer questions about them. We have no data on audio and video support, so it is better to choose a specialized model from the catalog for such files.
The model belongs to the multilingual Qwen line and is regularly used for Russian-language tasks. We do not have exact data on prompt and output languages, so please test it with your scenarios: in BotHub, this is available without VPN or foreign cards, with payment in rubles.