GPT OSS 120B
Neural Network
GPT OSS 120B is an open OpenAI language model with Mixture-of-Experts architecture for reasoning tasks and agent scenarios.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works GPT OSS 120B?
GPT OSS 120B is an open-weights OpenAI model based on the Mixture-of-Experts architecture: 117B parameters, with about 5.1B active per pass. The model is optimized for production: high loads, long chains of thought, and agent scenarios where tasks need to be broken down into steps. In BotHub, this model works from Russia without a VPN or foreign card, with payment in rubles via Russian cards for actual tokens used, which do not expire. Over 250 models are available in one window for comparison, and a unified OpenAI-compatible API connects the model to your service, allowing you to switch models without rewriting integrations; traffic is encrypted with AES-GCM. Developers can use it for code generation, debugging, and refactoring; analysts for step-by-step calculations and STEM tasks; product teams for building agents and internal assistants. Companies have access to contracts, invoices, electronic document management, and an admin panel with limits.Frequently asked questions about GPT OSS 120B
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
gpt-oss-120b is an open model from OpenAI for text processing: writing and editing code, analyzing documents, and solving reasoning tasks. It accepts images and documents as input and supports function calling. Suitable for assistants, analytics, and multimodal services via the unified BotHub API.
The output limit is 117,964 tokens per response, with a total context window of 131,072 tokens. This is sufficient for long documents, extensive code, and detailed analyses. Note that the request and response share the total context, so a large input reduces the space available for the response.
Yes, reasoning before answering is a feature of the model: it breaks the task into steps before providing a result. This significantly helps with math, logic, code debugging, and multi-step tasks where a verified answer is more important than a fast one.
Yes, the model supports function calling and structured JSON output. You can connect external tools, databases, and services, receiving responses in a predictable format. In BotHub, this works via a unified OpenAI-compatible API, so you can switch models without rewriting integrations.
You can upload images and documents: the model reads their content and responds with text. Audio and video are not listed as input formats in our data, nor is image or sound generation — for that, there are specialized models in the BotHub catalog available in the same interface.
We do not have data on prompt and response languages, so we cannot promise specific quality for Russian — it is easier to test it with your own task directly in BotHub. The service itself is fully in Russian: payment with Russian cards, access without VPN or foreign cards.