Video Generation
Wan Series
Alibaba Tongyi Wanxiang video models: text, image, reference and video editing on one task endpoint, billed by resolution tier × seconds
POST
Wan Series
Wan is Alibaba Tongyi’s video model line. This page covers nine video models:
A successful submission returns a
See Submit Task for
You are only billed for a successfully generated video. Failed and cancelled tasks, and tasks that return no usable video, are refunded in full. The final charge is the
Generated videos are returned in
wan2.5-preview, wan2.6, wan2.6-i2v and wan2.6-i2v-flash handle text- and image-to-video; wan2.7 and wan3.0-video / wan3.0-video-prime are multimodal entry points (text, frames, reference assets and video continuation all use one model ID); wan2.7-r2v does reference-to-video; wan2.7-videoedit does video editing. All nine share POST /v1/tasks and switching models means changing model, but the parameter set, the media inputs each takes, and the billing dimensions differ. For Wan image models, see Wan Image Series.
wan3.0-video-prime is the accelerated edition of wan3.0-video: the same parameters, value ranges, constraints and billing dimensions, with faster generation at a higher rate.
Quick start
task_id. Poll GET /v1/tasks/{task_id} for the result, wait inside the call with Prefer: wait, or configure a webhook.
Request parameters
string
required
One of nine:
wan2.5-preview, wan2.6, wan2.6-i2v, wan2.6-i2v-flash, wan2.7, wan2.7-r2v, wan2.7-videoedit, wan3.0-video / wan3.0-video-prime.string
default:"generate"
generate— generate a video. The only action for the seven models other thanwan2.7-videoeditedit— edit an input video. Onlywan2.7-videoeditsupports it, and it supports nothing else
string
Description of the video, or the edit instruction.Required for
wan2.7-r2v. Required for wan2.7 and wan3.0-video / wan3.0-video-prime when no media input is present. Optional elsewhere, but useful to guide camera and motion.string
Content to avoid.
wan3.0-video / wan3.0-video-prime does not take this parameter; the other seven models do.string[]
Image input. Meaning and caps differ by model:
wan2.5-preview/wan2.6— passing it switches to image-to-video modewan2.6-i2v/wan2.6-i2v-flash— required, exactly one image used as the first frame; the aspect ratio follows the imagewan2.7— 1–2 images (one = first frame, two = first and last frame); mutually exclusive withfirst_frame_image,video_urlsandaspect_ratiowan2.7-r2v— 1–5 subject / style references; mutually exclusive withvideo_urls; at most 4 when combined withfirst_frame_imagewan2.7-videoedit— up to 4 style / content referenceswan3.0-video/wan3.0-video-prime— up to 10; withgeneration_type: "frame"only 1–2 frame images
string
First-frame image URL. Taken by
wan2.7, wan2.7-r2v and wan3.0-video / wan3.0-video-prime.wan2.7— cannot be mixed withimage_urls,video_urlsoraspect_ratiowan2.7-r2v— may accompany up to 4 reference images; cannot be mixed withvideo_urlsoraspect_ratio, and the ratio follows the first framewan3.0-video/wan3.0-video-prime— cannot be mixed withimage_urls,video_urls,audio_urls,file_urlorlink_url
string
Last-frame image URL. Taken by
wan2.7 and wan3.0-video / wan3.0-video-prime, and only together with first_frame_image.string[]
Video input. Clip count and duration caps differ by model:
wan2.7— at most 1 clip, 2–10 s, for video continuationwan2.7-r2v— at most 5 clips, 1–30 s eachwan2.7-videoedit— required, exactly 1 clip, 2–10 swan3.0-video/wan3.0-video-prime— at most 5 clips, each up to 15 s, 15 s in total
string[]
Custom audio.
wan2.6-i2v-flash takes at most 1 clip (3–30 s), wan2.7 at most 1 clip (2–30 s), wan3.0-video / wan3.0-video-prime at most 5 clips (each 1–15 s, 15 s in total). The other models do not take it.boolean
Whether the output carries an audio track.
wan2.6 defaults to false; wan2.6-i2v-flash and wan3.0-video / wan3.0-video-prime default to true; wan2.5-preview only generates videos with audio, defaults to true and does not accept false. The other models do not take this parameter.string
Output resolution:
wan2.5-preview—480p,720p,1080pwan2.6,wan2.6-i2v,wan2.6-i2v-flash,wan2.7,wan2.7-r2v,wan2.7-videoedit—720p,1080pwan3.0-video/wan3.0-video-prime—480p,720p,1080p
480p for wan3.0-video / wan3.0-video-prime, 720p for the other models.integer
Output length in seconds:
wan2.5-preview—5or10wan2.6—5,10or15wan2.6-i2v,wan2.6-i2v-flash,wan2.7,wan2.7-r2v— any integer2–15wan2.7-videoedit— any integer2–10, plus0(keep the source video length)wan3.0-video/wan3.0-video-prime— any integer2–30; omit it or send0to let the model pick the length (2–30 seconds; with a reference video, input plus output at most 30 seconds), billed on the seconds actually produced
wan2.6-i2v defaults to 5. With video_urls present the output length is capped further: 10 seconds for wan2.7 and wan2.7-r2v, 15 seconds for wan3.0-video / wan3.0-video-prime.string
Aspect ratio.
wan2.5-preview, wan2.6, wan2.7, wan2.7-r2v and wan2.7-videoedit support 16:9, 9:16, 1:1, 4:3 and 3:4; wan3.0-video / wan3.0-video-prime adds adaptive and defaults to it.wan2.6-i2v and wan2.6-i2v-flash do not take this parameter — the aspect ratio follows the input image. It also cannot be sent by wan2.5-preview or wan2.6 alongside image_urls, by wan2.7 alongside first_frame_image / image_urls / video_urls, or by wan2.7-r2v alongside first_frame_image; in those cases the aspect ratio follows the input asset.integer
Random seed. The same request with the same seed produces similar results. All nine models take it.
boolean
Whether to add a watermark to the output. All nine models take it.
boolean
default:"true"
Smart prompt expansion, a clear improvement for short prompts.
wan3.0-video / wan3.0-video-prime does not take this parameter; the other seven models do.wan2.6-i2v and wan2.6-i2v-flash require prompt_extend to be true whenever shot_type is set.string
Shot type:
single (one continuous shot) or multi (multi-shot narrative). Only wan2.6, wan2.6-i2v and wan2.6-i2v-flash take it.string
Effect template. It needs one image only, and the
prompt is ignored.Values for wan2.6-i2v and wan2.6-i2v-flash: squish, rotation, poke, inflate, dissolve, carousel, singleheart, flying, rose, hug, frenchkiss, coupleheart. wan2.6 takes an effect template name as a string.string
Audio URL; takes precedence over
generate_audio. Only wan2.5-preview and wan2.6 take it.array
Role-tagged image array of
{url, role}, where role is first_frame, last_frame or reference; merged with first_frame_image / last_frame_image. Only wan2.7, wan2.7-r2v and wan3.0-video / wan3.0-video-prime take it.string
How the image array is classified:
frame (first / last frame) or reference (reference images). Only wan3.0-video / wan3.0-video-prime takes it; omitted, the classification follows the input assets.string
Reference document URL. Only
wan3.0-video / wan3.0-video-prime takes it, and it is mutually exclusive with link_url.string
Reference web page URL. Only
wan3.0-video / wan3.0-video-prime takes it, and it is mutually exclusive with file_url.object
Extra options object; currently only
audio_setting: auto (generate audio) or origin (keep the source video’s audio). Only wan2.7-videoedit takes it.callback_url, callback_events, Prefer: wait, Idempotency-Key, and the maximum-cost header.
None of the nine models support n or quality. Any other parameter returns 400 and is not billed.
Limits
Pricing
The whole series is billed by resolution tier × output seconds; an input video adds a charge on its input seconds.
Audio tier (
wan2.6-i2v-flash): enabling audio generation (generate_audio: true, the default on this model) or sending audio_urls bills at the audio tier.
Input-video seconds are read at submit time; a wan2.7-videoedit source whose length cannot be read returns 400 (input_video_duration_unknown) and is not billed. See Video Generation Overview for how the length is read.
wan3.0-video / wan3.0-video-prime freeze quota at submission (at 30 seconds when duration is 0 or omitted) and settle on the actual result when the task completes, refunding the difference.
Per-model rates are in
price_config on GET /v1/models, and in the console’s Model Market.What a single call actually cost is the cost field on the task response — an integer quota at 500,000 quota = 1 USD.cost field on the task response — an integer quota, not dollars.
Response
Completed task
result.videos; url is an array and expires_at is when the video link expires. See Query Task Status for the full field reference.
Available models
Generation (action: "generate")
Video editing (
action: "edit")