Gemma 4 31B Instruct
Neural Network
Gemma 4 31B Instruct is a Google DeepMind multimodal model with 256K context and function calling in the unified BotHub API.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Gemma 4 31B Instruct?
Gemma 4 31B Instruct is a dense Google DeepMind multimodal model with 30.7 billion parameters: it accepts text and images as input and outputs text. A 256,000-token context window allows you to work with voluminous documents, long dialogues, and large code fragments in their entirety, while the switchable reasoning mode helps where step-by-step task analysis is needed. Native function calling turns the model into an executive link for agent scenarios: it selects the tool and generates arguments itself. In BotHub, you work with Gemma 4 31B Instruct without a VPN or foreign card, paying with a Russian card for tokens used, not a subscription. Over 250 models are gathered in one window, so you can compare responses from neighboring models in a couple of clicks, and a unified OpenAI-compatible API allows you to switch models without rewriting the integration; correspondence is protected by AES-GCM encryption. Companies can connect via contract with invoicing, EDI, and an admin panel with limits. Suitable for analyzing documents and screenshots, code review and refactoring, extracting data from images, building agents with external tools, and RAG scenarios on long context.Frequently asked questions about Gemma 4 31B Instruct
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is a Google text model with a 65,536 token context. It analyzes long documents and describes images, helps with code—from refactoring to debugging, solves logic and STEM tasks, calls functions, and works in batch mode. You can switch to another model in BotHub without rewriting the integration.
The model outputs up to 16,384 tokens per response. This is enough for a voluminous article, a detailed technical analysis, or a large code module. If the material is larger, break the task into parts: the 65,536 token context allows you to keep the source documents in the dialogue entirely.
Yes, reasoning before answering is a stated capability of the model. It can structure the solution step-by-step before providing the result—this is noticeable in math, logic, and code analysis tasks. Ask to show the reasoning in the prompt if you need a visible chain of steps.
Yes, the model supports function calling and structured JSON output. You describe the tools, it selects the necessary one and returns arguments in the specified schema—this is convenient for agents, document parsing, and integrations. A unified OpenAI-compatible API is available in BotHub, so the connection is no different from what you are used to.
Images and documents—yes: you can send a screenshot, diagram, photo, or file and ask for analysis, data extraction, or converting a table to text. Audio and video input is not noted in our data, so for voiceover or video analysis, it is better to choose a specialized model in BotHub.
There is no separate note on language quality in our data, so we will not promise accuracy—test the model on your task right in the chat and compare it with others in the same window. The service itself works from Russia without a VPN or foreign card, with payment in rubles.