How it works Nemotron 3 Super 120B A12B?
Nemotron 3 Super 120B A12B is an open NVIDIA model with a hybrid Mamba-Transformer architecture and a sparse MoE scheme: out of 120 billion parameters, about 12 billion are activated per request, saving computation on complex multi-step tasks. It was created for multi-agent applications where multiple roles exchange intermediate results and maintain shared logic. In BotHub, you can connect to Nemotron 3 Super without a VPN or foreign card: pay with Russian cards in rubles based on actual usage; tokens do not expire. With over 250 models in one window, it's easy to compare answers and switch, while a unified OpenAI-compatible API allows changing models in code without rewriting integrations. Chats are encrypted via AES-GCM, data is not saved, and companies have access to contracts, invoices, electronic document management, and an admin panel with limits. In practice, the model is used for agent orchestration and tool chains, analytical tasks with long chains of thought, prototyping assistants via API, internal services with predictable costs, and research experiments with open models.Frequently asked questions about Nemotron 3 Super 120B A12B
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is an NVIDIA text model with a million-token context. It is suitable for working with large documents and codebases: refactoring, debugging, explaining logic, analyzing reports, and STEM tasks. It accepts text, documents, and images as input and supports function calling.
In a single response, the model generates up to 235,929 tokens — enough for voluminous technical text, a large code module, or a detailed analysis of an entire document. If the task is larger, break it down into sequential requests within the same dialogue.
Yes, the reasoning mode is confirmed: before the final answer, the model builds a chain of steps. This significantly helps with math, logic, code analysis, and multi-step tasks where it is important not just to guess the answer, but to derive it sequentially and verify intermediate conclusions.
Yes. The model supports function calling and structured JSON output, making it easy to integrate into agents and pipelines. In BotHub, it is available via a unified OpenAI-compatible API — you won't have to rewrite the integration when switching to another model.
You can upload images and documents — the model analyzes their content and responds with text. Audio and video input is not noted in our data, so do not count on them; for such formats, BotHub has separate specialized models.
We do not have data on prompt and response languages, so we will not promise specific quality — it is easier to test it with your own task in a few requests. The BotHub interface is in Russian, access is without VPN or foreign cards, and payment is in rubles based on actual usage.