Grok Imagine Video
Neural Network
Grok Imagine Video is an xAI model for generating videos from text descriptions, images, and references.
Max answer length
(in tokens)
Context size
(in tokens)
Video generation
(duration 8 sec., 720p)
Video generation
(duration 8 sec., 1080p)
How it works Grok Imagine Video?
Grok Imagine Video is a video generation model from xAI that works in three modes: text-to-video, image-to-video, and reference-based. It outputs short clips lasting 1 to 15 seconds at 24 fps, in 480p or 720p resolution, and in seven aspect ratios, making it suitable for both horizontal and vertical formats. The focus is on speed: you can regenerate a scene multiple times to refine movement and composition. In BotHub, Grok Imagine Video is accessible from Russia without a VPN or foreign card, with payment via Russian cards in rubles on a pay-as-you-go basis using non-expiring caps. Over 250 models are available in one window: create a storyboard with a text model, draw a reference frame, and immediately turn it into video. Product teams will benefit from a unified API, request encryption, and corporate access via contract. Use cases include promos for marketplaces, vertical social media clips, animating static illustrations and photos, scene pre-visualization, and dynamic backgrounds for presentations.Frequently asked questions about Grok Imagine Video
You can use generated results for commercial purposes. You own all rights to the content you create. The only restriction: make sure your prompt does not include copyrighted third-party material. You are responsible for respecting the rights to any input data.
grok-imagine-video generates videos from text descriptions or images. It is suitable for short clips, animating static frames, advertising concepts, scene previews, and visual mockups. In BotHub, the model is available from Russia without a VPN or foreign card, with payment via Russian cards.
The model works in two modes: text-to-video and image-to-video, outputting a finished video clip. We do not have fixed resolution and duration parameters in our data, so please refer to the technical specifications at the top of the page and the generation result.
A separate step-by-step reasoning mode is not noted in our data; this does not mean it doesn't exist, we just haven't confirmed it. For video generation, a clear prompt is more important: describe the scene, camera movement, and style, and the model will assemble the clip.
Support for function calling and strict JSON is not noted for this model in our data. However, you can access it via the unified OpenAI-compatible BotHub API, so integration into your service and switching to other models remains simple.
Images — yes: you can attach a frame or reference and animate it in image-to-video mode. Accepting audio and video as input is not noted in our data, so for such scenarios, it is better to check the result with a short test beforehand.
We do not have data on prompt languages for this model, so we cannot make any claims — please check with a short generation. However, BotHub itself is fully in Russian: the interface, website, and Telegram bot, with payment in rubles via Russian cards, without VPN or subscription.