Skip to main content
Generate high-quality AI videos using Kling O1. Supports text-to-video, raw reference images, persistent subject/object elements (image only), start/end frames, video transform, and video reference. Same shape as Omni 3.0 but does not support multi_shots, native_audio, or video-type elements. Max duration 10s.

Model

Parameters

Modes

Text-to-video with elements — set video_mode: "elements" and pass only the elements array (no image_urls required). The omit-video_mode path is strictly for pure text-to-video with zero inputs; any reference (image or element) requires elements mode.

Reference images vs elements

  • image_urls — raw images uploaded per request. Single-use, no description, no reuse across tasks. Cap 7. Reference in prompt as @Image1, @Image2, …
  • elements — persistent subject/object assets stored on the Kling account. Auto-created from your input on submit and auto-deleted if the submit fails. Reference in prompt as @Element1, @Element2, …
Both can be used together; they share the per-mode combined cap.

Elements

O1 supports IMAGE elements only — passing an element with type: "video" returns 400.
Each element accepts:

Example - Text-to-Video

Example - Reference Images

Example - With Elements

Example - Start/End Frame

Response