How it works Inkling Small?
Inkling Small is an open-weights multimodal model by Thinking Machines Lab using a mixture-of-experts architecture: out of 276 billion parameters, about 12 billion are activated per request, requiring fewer computations than a dense model of comparable size. In the Inkling line, this is the smaller, more economical version designed for streaming tasks where response speed and predictable processing costs are critical. In BotHub, the model is accessible without a VPN or foreign card: pay in rubles with a Russian card for tokens used, and unused balances do not expire. Access over 250 neural networks in one window, switch between them, and compare responses for the same prompt. A unified OpenAI-compatible API allows you to swap models in your system without rewriting integrations. Traffic is encrypted via AES-GCM, and chat content is not stored. The model is useful for mass text processing, multimodal queries, coding, extracting facts from documents, and prototyping assistants. Corporate clients can get contracts, invoices, EDI, and an admin panel with limits from BotHub.Frequently asked questions about Inkling Small
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
This is a text model with a reasoning mode: it breaks down tasks step-by-step, writes and edits code, explains errors, and accepts documents and images as input. The 262,144-token context window allows uploading large files and long conversations in full. Suitable for analytics, STEM tasks, and agent scenarios with tool calling.
The output limit is up to 262,144 tokens, the same as the context window capacity. This is enough for a large document, detailed analysis, or a large volume of code. Keep in mind that the request and response share the total context budget, so the practical available volume depends on the size of the prompt and attachments.
Yes, the model's reasoning mode is confirmed: it builds a chain of steps before the final answer. This helps with math, logic, code debugging, and multi-step tasks where a verifiable solution path is important. Note that reasoning consumes output tokens, so these responses are longer and take more time.
Yes, both function calling and structured JSON output are supported. The model can be connected to external tools, databases, and services to build agents and pipelines. In BotHub, this is available via a unified OpenAI-compatible API, so you don't need to rewrite integrations when switching to another model.
According to our data, text, images, and documents are accepted as input: send a screenshot, diagram, PDF, or table and get a text analysis. Audio and video are not listed among the supported input formats. The model is not noted for generating anything other than text — it responds in text.
Our data does not specify the language of requests and responses, so we cannot promise specific quality for Russian — test the model on your task in the chat or via API. There are no restrictions from the service side: access from Russia without a VPN or foreign card, payment in rubles.