Step 3.5 Flash
Neural Network
Step 3.5 Flash API is an open StepFun model based on Mixture of Experts architecture for text tasks and service integration.
Max answer length
(in tokens)
Context size
(in tokens)
Prompt cost
(per 1M tokens)
Answer cost
(per 1M tokens)
How it works Step 3.5 Flash?
Step 3.5 Flash is an open StepFun language model based on a sparse Mixture of Experts architecture: only 11 billion out of 196 billion parameters are activated per token, ensuring efficient computation with a vast knowledge base. In BotHub, it is available without VPN or foreign cards: pay with Russian cards in rubles only for tokens actually used, which do not expire. Over 250 neural networks are available in one window, allowing you to compare answers or switch models in a few clicks, while a unified OpenAI-compatible API eliminates the need to rewrite integrations. Dialogues are encrypted via AES-GCM, data is not saved, and for companies, we offer contracts, invoices, EDI, an admin panel, and employee limits. Use Step 3.5 Flash when you need to process user inquiries, draft and rewrite texts, summarize long documents, or test the model as a backup in a multi-model service.Frequently asked questions about Step 3.5 Flash
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
step-3.5-flash by stepfun is a text model with reasoning mode, image and document input support, and function calling. A 256,000-token context allows for analyzing large documents, codebases, and logs, writing and debugging code, preparing analytical summaries, and building agent scenarios.
The maximum volume of a single response is 65,536 tokens. This is enough for a large technical text, a detailed document analysis, or a large code fragment in its entirety, without splitting. If the response hits the limit, just ask it to continue — the model will finish in the next message.
Yes, reasoning before answering is a feature of the model: it goes through intermediate steps before providing a result. This is noticeable in math, logic, requirements analysis, and debugging — where it is important not just to guess the answer, but to follow the chain of thought.
Yes, both. Thanks to function calling, the model can be connected to search, databases, calculators, and internal services, while structured JSON output is convenient when the result goes directly into code without manual parsing. All via the unified OpenAI-compatible BotHub API.
According to our data, it accepts text, images, and documents as input: you can send a screenshot, diagram, table, or file and ask questions about the content. We have no specific notes on audio and video, so you should test these formats with your specific task.
We do not have specific data on language quality, so we won't make any promises — test it with your prompt, it's quick. The BotHub interface and Telegram bot are in Russian, access from Russia is available without VPN or foreign cards, and payment is via Russian cards.