Skip to main content
POST
Wan Series
Wan is Alibaba Tongyi’s video model line. This page covers nine video models: wan2.5-preview, wan2.6, wan2.6-i2v and wan2.6-i2v-flash handle text- and image-to-video; wan2.7 and wan3.0-video / wan3.0-video-prime are multimodal entry points (text, frames, reference assets and video continuation all use one model ID); wan2.7-r2v does reference-to-video; wan2.7-videoedit does video editing. All nine share POST /v1/tasks and switching models means changing model, but the parameter set, the media inputs each takes, and the billing dimensions differ. For Wan image models, see Wan Image Series. wan3.0-video-prime is the accelerated edition of wan3.0-video: the same parameters, value ranges, constraints and billing dimensions, with faster generation at a higher rate.

Quick start

A successful submission returns a task_id. Poll GET /v1/tasks/{task_id} for the result, wait inside the call with Prefer: wait, or configure a webhook.

Request parameters

string
required
One of nine: wan2.5-preview, wan2.6, wan2.6-i2v, wan2.6-i2v-flash, wan2.7, wan2.7-r2v, wan2.7-videoedit, wan3.0-video / wan3.0-video-prime.
string
default:"generate"
  • generate — generate a video. The only action for the seven models other than wan2.7-videoedit
  • edit — edit an input video. Only wan2.7-videoedit supports it, and it supports nothing else
string
Description of the video, or the edit instruction.Required for wan2.7-r2v. Required for wan2.7 and wan3.0-video / wan3.0-video-prime when no media input is present. Optional elsewhere, but useful to guide camera and motion.
string
Content to avoid. wan3.0-video / wan3.0-video-prime does not take this parameter; the other seven models do.
string[]
Image input. Meaning and caps differ by model:
  • wan2.5-preview / wan2.6 — passing it switches to image-to-video mode
  • wan2.6-i2v / wan2.6-i2v-flash — required, exactly one image used as the first frame; the aspect ratio follows the image
  • wan2.7 — 1–2 images (one = first frame, two = first and last frame); mutually exclusive with first_frame_image, video_urls and aspect_ratio
  • wan2.7-r2v — 1–5 subject / style references; mutually exclusive with video_urls; at most 4 when combined with first_frame_image
  • wan2.7-videoedit — up to 4 style / content references
  • wan3.0-video / wan3.0-video-prime — up to 10; with generation_type: "frame" only 1–2 frame images
string
First-frame image URL. Taken by wan2.7, wan2.7-r2v and wan3.0-video / wan3.0-video-prime.
  • wan2.7 — cannot be mixed with image_urls, video_urls or aspect_ratio
  • wan2.7-r2v — may accompany up to 4 reference images; cannot be mixed with video_urls or aspect_ratio, and the ratio follows the first frame
  • wan3.0-video / wan3.0-video-prime — cannot be mixed with image_urls, video_urls, audio_urls, file_url or link_url
string
Last-frame image URL. Taken by wan2.7 and wan3.0-video / wan3.0-video-prime, and only together with first_frame_image.
string[]
Video input. Clip count and duration caps differ by model:
  • wan2.7 — at most 1 clip, 2–10 s, for video continuation
  • wan2.7-r2v — at most 5 clips, 1–30 s each
  • wan2.7-videoedit — required, exactly 1 clip, 2–10 s
  • wan3.0-video / wan3.0-video-prime — at most 5 clips, each up to 15 s, 15 s in total
string[]
Custom audio. wan2.6-i2v-flash takes at most 1 clip (3–30 s), wan2.7 at most 1 clip (2–30 s), wan3.0-video / wan3.0-video-prime at most 5 clips (each 1–15 s, 15 s in total). The other models do not take it.
boolean
Whether the output carries an audio track. wan2.6 defaults to false; wan2.6-i2v-flash and wan3.0-video / wan3.0-video-prime default to true; wan2.5-preview only generates videos with audio, defaults to true and does not accept false. The other models do not take this parameter.
string
Output resolution:
  • wan2.5-preview — 480p, 720p, 1080p
  • wan2.6, wan2.6-i2v, wan2.6-i2v-flash, wan2.7, wan2.7-r2v, wan2.7-videoedit — 720p, 1080p
  • wan3.0-video / wan3.0-video-prime — 480p, 720p, 1080p
Defaults: 480p for wan3.0-video / wan3.0-video-prime, 720p for the other models.
integer
Output length in seconds:
  • wan2.5-preview — 5 or 10
  • wan2.6 — 5, 10 or 15
  • wan2.6-i2v, wan2.6-i2v-flash, wan2.7, wan2.7-r2v — any integer 2–15
  • wan2.7-videoedit — any integer 2–10, plus 0 (keep the source video length)
  • wan3.0-video / wan3.0-video-prime — any integer 2–30; omit it or send 0 to let the model pick the length (2–30 seconds; with a reference video, input plus output at most 30 seconds), billed on the seconds actually produced
wan2.6-i2v defaults to 5. With video_urls present the output length is capped further: 10 seconds for wan2.7 and wan2.7-r2v, 15 seconds for wan3.0-video / wan3.0-video-prime.
string
Aspect ratio. wan2.5-preview, wan2.6, wan2.7, wan2.7-r2v and wan2.7-videoedit support 16:9, 9:16, 1:1, 4:3 and 3:4; wan3.0-video / wan3.0-video-prime adds adaptive and defaults to it.wan2.6-i2v and wan2.6-i2v-flash do not take this parameter — the aspect ratio follows the input image. It also cannot be sent by wan2.5-preview or wan2.6 alongside image_urls, by wan2.7 alongside first_frame_image / image_urls / video_urls, or by wan2.7-r2v alongside first_frame_image; in those cases the aspect ratio follows the input asset.
integer
Random seed. The same request with the same seed produces similar results. All nine models take it.
boolean
Whether to add a watermark to the output. All nine models take it.
boolean
default:"true"
Smart prompt expansion, a clear improvement for short prompts. wan3.0-video / wan3.0-video-prime does not take this parameter; the other seven models do.wan2.6-i2v and wan2.6-i2v-flash require prompt_extend to be true whenever shot_type is set.
string
Shot type: single (one continuous shot) or multi (multi-shot narrative). Only wan2.6, wan2.6-i2v and wan2.6-i2v-flash take it.
string
Effect template. It needs one image only, and the prompt is ignored.Values for wan2.6-i2v and wan2.6-i2v-flash: squish, rotation, poke, inflate, dissolve, carousel, singleheart, flying, rose, hug, frenchkiss, coupleheart. wan2.6 takes an effect template name as a string.
string
Audio URL; takes precedence over generate_audio. Only wan2.5-preview and wan2.6 take it.
array
Role-tagged image array of {url, role}, where role is first_frame, last_frame or reference; merged with first_frame_image / last_frame_image. Only wan2.7, wan2.7-r2v and wan3.0-video / wan3.0-video-prime take it.
string
How the image array is classified: frame (first / last frame) or reference (reference images). Only wan3.0-video / wan3.0-video-prime takes it; omitted, the classification follows the input assets.
string
Reference document URL. Only wan3.0-video / wan3.0-video-prime takes it, and it is mutually exclusive with link_url.
Reference web page URL. Only wan3.0-video / wan3.0-video-prime takes it, and it is mutually exclusive with file_url.
object
Extra options object; currently only audio_setting: auto (generate audio) or origin (keep the source video’s audio). Only wan2.7-videoedit takes it.
See Submit Task for callback_url, callback_events, Prefer: wait, Idempotency-Key, and the maximum-cost header. None of the nine models support n or quality. Any other parameter returns 400 and is not billed.

Limits

Pricing

The whole series is billed by resolution tier × output seconds; an input video adds a charge on its input seconds. Audio tier (wan2.6-i2v-flash): enabling audio generation (generate_audio: true, the default on this model) or sending audio_urls bills at the audio tier. Input-video seconds are read at submit time; a wan2.7-videoedit source whose length cannot be read returns 400 (input_video_duration_unknown) and is not billed. See Video Generation Overview for how the length is read. wan3.0-video / wan3.0-video-prime freeze quota at submission (at 30 seconds when duration is 0 or omitted) and settle on the actual result when the task completes, refunding the difference.
Per-model rates are in price_config on GET /v1/models, and in the console’s Model Market.What a single call actually cost is the cost field on the task response — an integer quota at 500,000 quota = 1 USD.
You are only billed for a successfully generated video. Failed and cancelled tasks, and tasks that return no usable video, are refunded in full. The final charge is the cost field on the task response — an integer quota, not dollars.

Response

Completed task
Generated videos are returned in result.videos; url is an array and expires_at is when the video link expires. See Query Task Status for the full field reference.

Available models

Generation (action: "generate") Video editing (action: "edit")