Video Generation
Kling Series
Kuaishou Kling video models: text-to-video, first/last frames, image and video references, video editing and motion control, 4K on kling-v3 / kling-v3-omni
POST
Kling Series
Kuaishou’s Kling video line currently has seven models:
A successful submission returns a
Multi-shot and reference subjects (
Role-tagged images (
Motion control (
See Submit a task for
You are only billed for a successfully generated video. Failed and cancelled tasks, and tasks that return no usable video, are refunded in full. The final charge is the
Generated videos are returned in
kling-3.0-turbo, kling-v2-6 and kling-v3 for general generation, kling-v3-omni and kling-video-o1 for video references and video editing, and kling-v2-6-motion-control and kling-v3-motion-control for motion transfer. They share the task endpoint POST /v1/tasks — switching models means changing model — but their actions, parameters and durations differ; see Available models for the per-model breakdown.
Quick start
task_id. Retrieve the result with GET /v1/tasks/{task_id}, wait inline with Prefer: wait, or configure a webhook.
Request parameters
string
required
Model ID; see Available models.
string
default:"generate"
generate— generate a video; supported by all seven modelsedit— edit a source video; onlykling-v3-omniandkling-video-o1support it, andvideo_urlsis required
string
required
Video description or edit instruction. On
kling-v3-omni and kling-video-o1, <<<image_N>>> references the Nth entry of image_urls (N starts at 1); a prompt without any reference gets <<<image_1>>> prepended. with element_list on kling-v3 and kling-v3-omni, reference a subject as @name.integer
default:"1"
Number of videos. Value:
1. Each task returns one video.integer
required
Output duration in seconds, per model:
kling-3.0-turbo,kling-v3,kling-v3-omni— an integer from3to15kling-v2-6,kling-video-o1—5or10kling-v2-6-motion-control,kling-v3-motion-control—0, meaning the reference video’s lengthaction: "edit"onkling-v3-omni/kling-video-o1—0, meaning the source video’s length
string
default:"720p"
Output resolution tier:
720p or 1080p; kling-v3 and kling-v3-omni also offer 4k.string
Frame ratio:
16:9, 9:16 or 1:1. The two *-motion-control models do not take this parameter. With reference images, the output follows the input image’s ratio.string[]
Reference image URLs:
kling-3.0-turbo— up to 1, used as the first framekling-v2-6,kling-v3,kling-video-o1— up to 2; two images act as first and last framekling-v3-omni— up to 7, together with multi-image subjects inelement_listno more than 7- The two
*-motion-controlmodels — required, one character reference image - With
video_urlsonkling-v3-omni/kling-video-o1— at most 1;action: "edit"takes no reference image
string[]
Reference video URLs, at most one entry:
kling-v3-omni,kling-video-o1— a camera/feature reference underaction: "generate", the source clip underaction: "edit"; 3 to 10 seconds- The two
*-motion-controlmodels — required, the motion source clip; up to 30 seconds
string
Negative prompt. Supported by
kling-v2-6, kling-v3 and kling-v3-omni.boolean
default:"false"
Generate native audio. Supported by
kling-v2-6, kling-v3 and kling-v3-omni; on kling-v2-6 it requires resolution: "1080p", and on kling-v3-omni it cannot be combined with video_urls.boolean
Whether to watermark the output video. Supported by every model except
kling-video-o1.Multi-shot and reference subjects (kling-v3 / kling-v3-omni)
boolean
default:"false"
Enable multi-shot storyboard mode.
string
Shot split method:
customize / intelligence; required when multi_shot=true.array
Multi-shot list (1-6 items) of
{index, prompt, duration}; index starts at 1 and is contiguous; shot durations must sum to duration; requires multi_shot=true.array
Reference subjects (up to 3) of
{name, description, element_input_urls}; 2-4 images per subject, first one frontal; reference a subject in the prompt as @name.Role-tagged images (kling-v3-omni / kling-video-o1)
array
Role-tagged image array of
{url, role} (first_frame / last_frame / reference); merged with first_frame_image / last_frame_image.Motion control (kling-v2-6-motion-control / kling-v3-motion-control)
string
required
Use the reference image or video for character orientation:
image or video. The two values imply different reference-video lengths; see Limits.string
default:"yes"
Whether to keep the motion reference video’s original audio:
yes or no.callback_url, callback_events, Prefer: wait, Idempotency-Key, and the maximum-cost header.
The Kling series does not support seed, quality or size. Any other parameter returns 400 and is not billed.
Limits
Requests whose duration follows a reference video (both
*-motion-control models, and action: "edit") need a video link whose duration can be read. If it cannot, the request returns 400 and is not billed; ordinary MP4 and WebM links work.
Pricing
Billing is resolution tier × output seconds.resolution picks the tier (720p, 1080p or 4k), and the seconds are the clip’s actual output length. The number of reference images does not affect the price.
Two situations switch to a different per-second rate while still multiplying by output seconds only:
- Native audio —
generate_audio: trueuses the audio tier.kling-v3andkling-v3-omnihave one at both resolutions;kling-v2-6has one at 1080p only. - Video reference — on
kling-v3-omniandkling-video-o1, supplyingvideo_urlsuses the video-reference tier, under bothaction: "generate"andaction: "edit". The reference video’s own length is not billed.
4k tier (kling-v3, kling-v3-omni) has a single per-second rate; native audio and a reference video do not change it.
Where the output seconds come from:
- Requests that carry a concrete
durationare billed on thatduration. - Requests with
duration: 0(both*-motion-controlmodels, andaction: "edit") follow the reference or source video, and are billed on the duration read from that video.
Per-model rates are in
price_config on GET /v1/models, and in the console’s Model Market.What a single call actually cost is the cost field on the task response — an integer quota at 500,000 quota = 1 USD.cost field on the task response — an integer quota, not dollars.
Response
Completed task
result.videos; expires_at is when the link stops working, so store the file before then. See Query Task Status for the full field list.
Available models
Picking by capability:
- Native audio (
generate_audio) —kling-v2-6(1080ponly),kling-v3,kling-v3-omni - Multi-shot storyboards and reference subjects (
multi_shot/shot_type/multi_prompt/element_list) —kling-v3,kling-v3-omni - Role-tagged images (
image_with_roles) —kling-v3-omni,kling-video-o1 - Video editing and video references (
action: "edit"/video_urls) —kling-v3-omni,kling-video-o1 - Motion transfer (
character_orientation/keep_original_sound) —kling-v2-6-motion-control,kling-v3-motion-control - Negative prompts (
negative_prompt) —kling-v2-6,kling-v3,kling-v3-omni