Gemma 3 4B IT
Neural Network
Gemma 3 4B IT is a multimodal Google model for text and images; Gemma 3 API with context up to 128k tokens.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Gemma 3 4B IT?
Gemma 3 4B IT is a multimodal Google model: it accepts text and images as input and responds with text. A context of up to 128k tokens allows you to keep a long document or a large code snippet in a single request. The model understands over 140 languages, and compared to the previous Gemma generation, it has improved math, reasoning, and dialogue quality. In BotHub, the model is accessible without a VPN or foreign card: you pay with a Russian card in rubles, only for the tokens you use, without a subscription, and tokens do not expire. Over 250 neural networks are available in one window, making it easy to switch between them and compare answers. A unified OpenAI-compatible API allows you to connect Gemma 3 to your service and replace it with another model without rewriting the integration. Chats are encrypted via AES-GCM, and companies have access to contracts, invoices, electronic document management (EDO), and an admin panel with limits. Developers can use the model for code analysis and error explanation, analysts for summaries of long documents and tables, support teams for classifying tickets and drafting responses, and those working with visuals to get text descriptions of screenshots, diagrams, or photos.Frequently asked questions about Gemma 3 4B IT
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is a compact Google model from the Gemma family. It accepts text, documents, and images, and responds with text: suitable for drafting and rewriting, summarizing and structuring, answering questions about uploaded files, analyzing screenshots, and simple coding tasks. Context is 131,072 tokens.
Up to 16,384 tokens in a single response — enough for a long article or a large code snippet. If there is more material, just ask it to continue: the 131,072-token context allows it to keep previous parts of the dialogue and maintain the style.
Our data does not indicate a separate reasoning mode for this model, so the answer usually comes immediately. You can ask it to break down a task step-by-step directly in the prompt. If you need a specific reasoning model, switch to another one in the same window.
Our data does not indicate function calling or structured JSON output for this model. You can ask for a JSON response in text, but the result should be validated. For reliable tools, use the unified OpenAI-compatible BotHub API and connect a suitable model — you won't have to rewrite the integration.
Images and documents — yes: attach a photo, screenshot, or file and ask it to describe the content, extract data, or explain an error on the screen. Audio and video input is not indicated in our data — BotHub has separate models for sound and video.
We do not have specific data on supported languages, so we won't make any specific promises — it's faster to test the model with your typical queries. BotHub itself works from Russia without a VPN or foreign card: the interface is in Russian, and payment is made with Russian cards in rubles.