Video Generation
PixVerse V6
PixVerse V6: text-to-video, image-to-video, first/last frame, multi-reference fusion and video extension, any integer duration from 1 to 15 seconds, optional native audio
POST
PixVerse V6
PixVerse V6 is PixVerse’s unified video model: one model covers text-to-video, image-to-video, first/last-frame transitions and multi-image reference fusion, with a multi-shot engine for continuous narrative, native synchronized audio (dialogue, sound effects and background music) and cinematic camera control, at 360p to 1080p and 1 to 15 seconds. Under
A successful submission returns a
See Submit Task for
You are billed only for a delivered video. Failed and cancelled tasks, and upstream successes that carry no deliverable video, are refunded in full. The final charge is the
Generated videos are returned in
generate the mode follows the media fields you send; extend continues a completed task.
Quick start
task_id. Retrieve the result with GET /v1/tasks/{task_id}, wait inline with Prefer: wait, or configure a webhook.
Chaining an extension takes two steps. Submit action: "generate" and keep the task_id; once the task is completed, pass that task_id as ref_task_id on an action: "extend" request. ref_task_id is the task_id of a completed pixverse-v6 task in your account, not an upstream ID.
Request parameters
string
required
Must be
pixverse-v6.string
default:"generate"
generate— generate a video.promptalone is text-to-video,image_urlsis image-to-video,first_frame_imagetogether withlast_frame_imageis a first/last-frame transition, andimg_referencesis multi-image reference fusion.extend— video extension;ref_task_idis required.
string
required
The video description, up to 5000 characters. Under
action: "extend" it describes the continuation.string
Negative prompt used to exclude unwanted content, up to 2048 characters.
integer
Video length in seconds, an integer from
1 to 15. First/last-frame mode accepts only 5 or 8.string
default:"540p"
Output resolution:
360p, 540p, 720p or 1080p; defaults to 540p. Resolution selects the price tier.string
Frame ratio:
16:9, 4:3, 1:1, 3:4, 9:16, 2:3, 3:2 or 21:9. Effective in text-to-video and multi-reference fusion only; other modes follow the input image, and action: "extend" follows the source video.integer
Random seed from
0 to 2147483647. The same prompt and seed reproduce similar results.string[]
Input image URLs for image-to-video; only the first image is used.
action: "generate" only.string
Starting frame URL for a first/last-frame transition. Must be sent together with
last_frame_image. action: "generate" only.string
Closing frame URL for a first/last-frame transition. Must be sent together with
first_frame_image. action: "generate" only.boolean
default:"false"
Whether to add a watermark in the bottom-right corner of the video.
boolean
default:"false"
Whether to generate a video with an audio track. Enabling it moves the task to the audio price tier.
string
Motion mode.
pixverse-v6 supports normal only.boolean
default:"false"
Whether to generate a multi-clip continuous video. Effective in the text-to-video and image-to-video modes of
action: "generate" only.string[]
Reference image URLs for multi-image fusion, 1 to 7 images. Sending this field selects fusion mode.
action: "generate" only.string
The task to extend; required under
action: "extend". Use the task_id of a completed pixverse-v6 task in your account.callback_url, callback_events, Prefer: wait, Idempotency-Key and the maximum-cost header.
Any other parameter returns 400 and is not billed.
Limits
All image inputs must be publicly reachable HTTP(S) URLs; base64 and Data URIs return an error.
Pricing
Billing is resolution tier × output seconds, where the seconds come from theduration you send; without resolution the 540p tier applies. Each of the four resolutions is priced separately, and higher resolutions cost more. action: "extend" is billed the same way, on this request’s duration and resolution, independent of the source task.
generate_audio is the second billing dimension: enabling it settles at the audio tier for the same resolution, which is above the silent tier. Reference images, first/last frames, the negative prompt and the aspect ratio carry no separate charge.
Per-model rates are in
price_config on GET /v1/models, and in the console’s Model Market.What a single call actually cost is the cost field on the task response — an integer quota at 500,000 quota = 1 USD.cost field on the task response — an integer quota, not dollars (500,000 quota = 1 USD).
Response
Completed task
result.videos; url is an array and expires_at is the link expiry in Unix seconds, so copy the file to your own storage before then. See Query Task Status for the full field reference.