xAI upgraded the model behind
xai/grok-imagine-video: it used to run Grok Imagine 1.0 Video, and now runs a speed-optimized build of Grok Imagine 1.5 Video — faster results in exchange for a small amount of quality. This build differs from the official API version, which is available separately as Grok Imagine Video 1.5.Model
Parameters
Audio generation works for both text-to-video and requests with visual references. Set
audio: false only when you want to suppress generated audio. A preset voice still requires a visual input and cannot be combined with audio: false.Media reference objects
Transient
character, prop, and location references are created for the task and automatically deleted after the generation completes or fails. A caller-provided character_id is never deleted.
Video presets:
Example requests
Text to video
Text-only requests can generate audio. This example disables it explicitly.Image to video with generated audio
Reference images
Character with a preset voice
GET /v1/resources/grok/voices to discover all preset IDs, descriptions, tags, and preview URLs. Voice names are converted to Grok’s internal asset IDs by the API.
Prompt references
Prompt references are only available on
xai/grok-imagine-video. Grok Imagine Video 1.5 takes a single start_image_url and does not support reference assets.Response
Volume discounts are available for almost all models. Reach out at contact@unifically.com to discuss custom pricing.
