Skip to main content
Generate high-quality AI videos using xAI’s Grok Imagine model.
xAI upgraded the model behind xai/grok-imagine-video: it used to run Grok Imagine 1.0 Video, and now runs a speed-optimized build of Grok Imagine 1.5 Video — faster results in exchange for a small amount of quality. This build differs from the official API version, which is available separately as Grok Imagine Video 1.5.

Model

Parameters

Audio generation works for both text-to-video and requests with visual references. Set audio: false only when you want to suppress generated audio. A preset voice still requires a visual input and cannot be combined with audio: false.

Media reference objects

Transient character, prop, and location references are created for the task and automatically deleted after the generation completes or fails. A caller-provided character_id is never deleted. Video presets:

Example requests

Text to video

Text-only requests can generate audio. This example disables it explicitly.

Image to video with generated audio

Reference images

Character with a preset voice

Use GET /v1/resources/grok/voices to discover all preset IDs, descriptions, tags, and preview URLs. Voice names are converted to Grok’s internal asset IDs by the API.

Prompt references

Prompt references are only available on xai/grok-imagine-video. Grok Imagine Video 1.5 takes a single start_image_url and does not support reference assets.

Response

Volume discounts are available for almost all models. Reach out at contact@unifically.com to discuss custom pricing.