Skip to main content
QWave API supports many video generation models, all called through the unified async task endpoint /v1/tasks. Video generation typically takes from 10 seconds up to several minutes; after submission you receive a task_id, and you fetch the result via polling or a webhook callback.

Request Endpoint

For the full async flow see Task System — Submit → Status → Result + optional webhook.

Supported Models

Veo Series

Google Veo 3.1 — official Fast / Quality billed per second; economy Lite plus the Fast / Quality reverse tiers at a per-clip price, with extend on the reverse tiers

Seedance Series

ByteDance Seedance 1.0 Pro / 1.0 Pro Fast / 1.5 Pro / 2.0 / 2.0 Fast / 2.0 Mini / 2.5, with video edit and extend on 2.5

Kling Series

Kuaishou Kling 3.0 Turbo / v2.6 / v3 / v3 Omni / Video O1, plus two motion-control tiers

Hailuo Series

MiniMax Hailuo 02 / 2.3 / 2.3 Fast — text, first frame and first-and-last frame

MiniMax H3

MiniMax H3 / H3 Max text- and reference-to-video, with H3 Regeneration re-rendering a finished clip at native 2K

Vidu Q3 Series

Shengshu Vidu Q3 / Q3 Mix / Q3 Pro / Q3 Turbo — start frame, first-and-last frame and reference images

Vidu Q4 Preview

Shengshu Vidu Q4 Preview — start-frame video, or up to 15 reference images plus 3 reference audio clips, up to 4K

Wan Series

Alibaba Wan 2.5 / 2.6 / 2.6 i2v / 2.7 / 2.7 R2V / 2.7 VideoEdit / 3.0 — text, image, reference and video editing

HappyHorse Series

Alibaba Cloud HappyHorse 1.0 (with video editing) / 1.1 — text, image and reference video

FLUX 3 Video

Black Forest Labs FLUX 3 Video standard / Draft tiers, keyframes and continuation, with synchronized audio

SkyReels V4 Series

Kunlun SkyReels V4 Fast / Standard, multimodal reference

Grok Imagine Video Series

xAI Grok Imagine Video official, 1.5 official and reverse tiers — text- and image-to-video

PixVerse V6

PixVerse V6, optional audio generation

Omni Video Series

Omni-Flash-Ext at fixed lengths (first frame, reference images, reference video as motion reference), and Gemini Omni 1.1 Flash / Flash Preview with a model-chosen length

Common Parameters (shared across the series)

Each model’s actual supported range differs — see the individual vendor docs.
string
required
Model ID (group_name); pick from the vendor docs above
string
default:"generate"
generate — the only action this model takes.The mode is decided by what you send, not by the action: prompt alone is text-to-video; image_urls is image-to-video; first_frame_image + last_frame_image is a first/last-frame transition; video_urls / audio_urls is reference-to-video.
string
Video description text. Required for T2V; optional as guidance for other modes
integer
Video duration in seconds; range depends on the model
string
Frame aspect ratio, e.g. 16:9 / 9:16 / 1:1
string
Output resolution, e.g. 720p / 1080p / 4K
string[]
Array of reference image URLs (used for I2V / R2V)
string
First-frame image URL (used for first/last-frame mode)
string
Last-frame image URL; must be paired with first_frame_image
string[]
Array of reference video URLs — a single element is sufficient (used for video continuation / R2V video reference / video editing)
Reference-video billing differs by vendor — see each model page, and Input video duration and billing below.
string[]
Array of reference audio URLs — a single element is sufficient (used for driving audio / custom voiceover)
string
Webhook callback URL, invoked when the task reaches a terminal state. See Task System

Input Video Duration and Billing

Requests that carry video_urls (reference video / video editing / motion control) follow these three rules. How seconds are counted — the gateway reads the video’s file header for its real duration and truncates to whole seconds (10.9 seconds counts as 10); anything under 1 second counts as 1. This matches the upstream’s own accounting. What happens if it cannot be read — if the header cannot be parsed (unusual format, unreachable URL, and so on) the request is rejected with 400 input_video_duration_unknown and nothing is charged — it is not estimated at a cap and billed. Use a directly downloadable URL in a common container format (mp4 / webm). Estimating before you submit — POST https://api.qingbo.ai/waveapi/quote with video_urls returns the exact price for that call, including the input seconds it probed. This is a public utility endpoint and needs no API key. It replies in the standard envelope {code, message, data}, where data carries: When a request is rejected, errorKey carries the specific reason (such as input_video_duration_unknown) and message holds text you can show to the user directly. How the reference video then enters the price (counted as seconds, only switching tier, or replacing the clip price) depends on the vendor — see the billing section on each model page.

Submit Response Example

After receiving task_id, poll GET /v1/tasks/{task_id} until status = completed, then read the video URL.

Mode Quick Reference

Not every model supports every mode — check each vendor’s doc for the actual supported action list and field range. The backend validates whether request fields fall within the vendor’s declared capabilities and returns an error if they do not.

Field Naming Conventions

  • Media references are always plural — always image_urls / video_urls / audio_urls; even a single video uses a one-element array ["one.mp4"]
  • Frame ratio is unified as aspect_ratio — vendor-internal size / ratio etc. are implementation details and not exposed
  • Resolution is unified as resolution — vendor-internal mode / quality etc. are implementation details
  • First/last-frame fields carry the _image suffix — first_frame_image / last_frame_image (emphasizing the image resource)