Gemma 4 26B A4B IT
Neural Network
Gemma 4 26B A4B IT is Google DeepMind's instruction-tuned MoE model for text tasks and integrations via the Gemma 4 API.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Gemma 4 26B A4B IT?
Gemma 4 26B A4B IT is an instruction-tuned Google DeepMind model with a Mixture-of-Experts architecture: out of 25.2B parameters, only 3.8B are activated per token, maintaining quality comparable to dense ~31B models while requiring fewer computations. Instruction fine-tuning makes it predictable in dialogue and specific tasks. Via BotHub, the model works from Russia without a VPN or foreign card: you pay with a Russian card in rubles only for used tokens, with no subscription. Over 250 models are available in one window, making it easy to compare answers, and a unified OpenAI-compatible API allows switching models without rewriting integrations. Chats are encrypted via AES-GCM, and for companies, we offer contracts, invoices, EDI, and an admin panel with limits. Gemma 4 is typically used for assistants and chatbots with high request volumes, drafting and text processing, document analysis and summarization, and background product tasks where cost per response is critical.Frequently asked questions about Gemma 4 26B A4B IT
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is a Google text model with a 65,536 token context. It is suitable for coding, document and image analysis, analytics, step-by-step reasoning, and function calling integrations. Available on BotHub without VPN or foreign cards, with payment in rubles.
The output limit is up to 235,929 tokens per response, which is a very large volume: a long article, detailed analysis, or a large code snippet. However, the context window is 65,536 tokens, so plan your dialogue history and attachments accordingly.
Yes, reasoning before answering is confirmed. The model can break down a task into steps before providing a result. This significantly helps with logic, math, code debugging, and document analysis, where the process is as important as the final result.
Yes, both function calling and structured JSON output are supported. This makes it easy to connect the model to external tools, databases, and services. On BotHub, everything works via a unified OpenAI-compatible API, so you can switch models without rewriting your integration.
Images and documents, yes; you can send them with text and ask to analyze a screenshot, diagram, or file. Audio and video are not listed in our data, so for audio processing or video work, it is better to choose a specialized model from the BotHub catalog.
We do not have data on prompt and response languages, so we cannot make any promises—please test the model with your specific task. On BotHub, you pay only for the tokens used, and they do not expire, so you can easily compare several models.