Skip to main content
Generate videos using ByteDance SeeDance 2.5 with text-to-video, first/last frame, and multimodal reference workflows. It supports image, video, and audio references, synchronized native audio, 4–30 second durations, and 480p or 720p output.

Model

Unifically automatically routes prompt-only requests through the text generation channel and requests containing any reference media through the multimodal channel. No routing option is required in your request.

Request types

Parameters

The total number of reference assets across images, videos, and audio cannot exceed 50. First/last frame URLs cannot be combined with image_urls, video_urls, or audio_urls.

Example - Text-to-Video

Example - First & Last Frame

Example - Video and Audio References

References are numbered independently and in array order: the first image is [Image1], the first video is [Video1], and the first audio clip is [Audio1].

Response