Skip to main content
POST
SkyReels V4 Series
SkyReels V4 is a unified multimodal video-and-audio generation model that jointly produces synchronized picture and sound from text, image and video inputs, at 480p to 1080p and 3 to 15 seconds, and can extend an existing video. This page covers skyreels-v4-fast and skyreels-v4-std: both tiers expose the same actions, parameters and value ranges — Fast favors speed, Std favors fidelity — so switching tiers means changing model.

Quick start

A successful submission returns a task_id. Retrieve the result with GET /v1/tasks/{task_id}, wait inline with Prefer: wait, or configure a webhook.

Request parameters

string
required
Either skyreels-v4-fast or skyreels-v4-std.
string
default:"generate"
  • generate — generate a video. The mode follows the media fields you send: prompt alone is text-to-video; first_frame_image, end_frame_image or mid_frame_images is image-to-video; ref_images or video_urls is reference-driven generation. Image-to-video fields and reference fields cannot be combined.
  • extend — video extension. Continues the one source video in video_urls for duration seconds.
string
required
The video description. When you send video_urls, ref_images or mid_frame_images, the prompt must contain the matching @ tag; under action: "extend" reference @video_1.
integer
Video length in seconds, an integer from 3 to 15.
  • action: "generate" with video_urls: the output follows the reference video (3 seconds when shorter) and this parameter has no effect
  • action: "extend": the length of the extension
string
default:"1080p"
Output resolution: 480p, 720p or 1080p; defaults to 1080p.
string
Frame ratio: 16:9, 9:16, 1:1, 4:3 or 3:4. Ignored when a first/end frame image or a reference video is present, and under action: "extend" — the ratio follows the input material.
string[]
Video URLs, at most 1, MP4 / MOV. The Nth entry gets the tag @video_N, which the prompt must reference (for example Follow the camera motion in @video_1); without it the upstream rejects the request and nothing is billed.
  • action: "generate": motion and subject reference, 1 to 10 seconds
  • action: "extend": required, the source video to extend, 1 to 15 seconds
string
Starting frame image URL (jpg / jpeg / png / gif / bmp). action: "generate" only.
string
Closing frame image URL; combine with the first frame for start/end control. action: "generate" only.
object[]
Mid keyframe list, up to 6 entries, available in the image-to-video mode of action: "generate" only. Each entry holds:
  • tag: a tag starting with @ that must appear in the prompt
  • image_url: the image URL
  • time_stamp: the moment it appears, in seconds. -1 (the default) leaves it unspecified; otherwise 0 < time_stamp < duration
object[]
Reference image list, available in the reference mode of action: "generate" only and mutually exclusive with the image-to-video fields. All entries must share the same type. Each entry holds:
  • tag: a tag starting with @ that must appear in the prompt
  • type: image for an ordinary reference image, grid for a single image composed of several tiles
  • image_urls: an array of image URLs
  • audio_url: a voice audio URL, allowed only when type is image, up to 15 seconds
boolean
default:"true"
Whether to auto-optimize the prompt.
See Submit Task for callback_url, callback_events, Prefer: wait, Idempotency-Key and the maximum-cost header. Any other parameter returns 400 and is not billed.

Limits

All input material must be publicly downloadable over HTTP(S).

Pricing

Billing is output resolution tier × seconds; without resolution the 1080p tier applies. Every skyreels-v4-std tier sits above the matching skyreels-v4-fast tier. The reference-video tier is above the ordinary tier at the same resolution. The input video’s own seconds are not billed; reference images (ref_images), keyframe images and prompt length carry no separate charge.
Per-model rates are in price_config on GET /v1/models, and in the console’s Model Market.What a single call actually cost is the cost field on the task response — an integer quota at 500,000 quota = 1 USD.
You are billed only for a delivered video. Failed and cancelled tasks, and upstream successes that carry no deliverable video, are refunded in full. The final charge is the cost field on the task response — an integer quota, not dollars (500,000 quota = 1 USD).

Response

Completed task
Generated videos are returned in result.videos, with picture and audio in the same file; url is an array and expires_at is the link expiry in Unix seconds, so copy the file to your own storage before then. See Query Task Status for the full field reference.

Available models