/v1/tasks. Video generation typically takes from 10 seconds up to several minutes; after submission you receive a task_id, and you fetch the result via polling or a webhook callback.
Request Endpoint
Supported Models
Veo Series
Google Veo 3.1 — official Fast / Quality billed per second; economy Lite plus the Fast / Quality reverse tiers at a per-clip price, with extend on the reverse tiers
Seedance Series
ByteDance Seedance 1.0 Pro / 1.0 Pro Fast / 1.5 Pro / 2.0 / 2.0 Fast / 2.0 Mini / 2.5, with video edit and extend on 2.5
Kling Series
Kuaishou Kling 3.0 Turbo / v2.6 / v3 / v3 Omni / Video O1, plus two motion-control tiers
Hailuo Series
MiniMax Hailuo 02 / 2.3 / 2.3 Fast — text, first frame and first-and-last frame
MiniMax H3
MiniMax H3 / H3 Max text- and reference-to-video, with H3 Regeneration re-rendering a finished clip at native 2K
Vidu Q3 Series
Shengshu Vidu Q3 / Q3 Mix / Q3 Pro / Q3 Turbo — start frame, first-and-last frame and reference images
Vidu Q4 Preview
Shengshu Vidu Q4 Preview — start-frame video, or up to 15 reference images plus 3 reference audio clips, up to 4K
Wan Series
Alibaba Wan 2.5 / 2.6 / 2.6 i2v / 2.7 / 2.7 R2V / 2.7 VideoEdit / 3.0 — text, image, reference and video editing
HappyHorse Series
Alibaba Cloud HappyHorse 1.0 (with video editing) / 1.1 — text, image and reference video
FLUX 3 Video
Black Forest Labs FLUX 3 Video standard / Draft tiers, keyframes and continuation, with synchronized audio
SkyReels V4 Series
Kunlun SkyReels V4 Fast / Standard, multimodal reference
Grok Imagine Video Series
xAI Grok Imagine Video official, 1.5 official and reverse tiers — text- and image-to-video
PixVerse V6
PixVerse V6, optional audio generation
Omni Video Series
Omni-Flash-Ext at fixed lengths (first frame, reference images, reference video as motion reference), and Gemini Omni 1.1 Flash / Flash Preview with a model-chosen length
Common Parameters (shared across the series)
Each model’s actual supported range differs — see the individual vendor docs.
string
required
Model ID (group_name); pick from the vendor docs above
string
default:"generate"
generate — the only action this model takes.The mode is decided by what you send, not by the action: prompt alone is text-to-video; image_urls is image-to-video; first_frame_image + last_frame_image is a first/last-frame transition; video_urls / audio_urls is reference-to-video.string
Video description text. Required for T2V; optional as guidance for other modes
integer
Video duration in seconds; range depends on the model
string
Frame aspect ratio, e.g.
16:9 / 9:16 / 1:1string
Output resolution, e.g.
720p / 1080p / 4Kstring[]
Array of reference image URLs (used for I2V / R2V)
string
First-frame image URL (used for first/last-frame mode)
string
Last-frame image URL; must be paired with
first_frame_imagestring[]
Array of reference video URLs — a single element is sufficient (used for video continuation / R2V video reference / video editing)
Reference-video billing differs by vendor — see each model page, and Input video duration and billing below.
string[]
Array of reference audio URLs — a single element is sufficient (used for driving audio / custom voiceover)
string
Webhook callback URL, invoked when the task reaches a terminal state. See Task System
Input Video Duration and Billing
Requests that carryvideo_urls (reference video / video editing / motion control) follow these three rules.
How seconds are counted — the gateway reads the video’s file header for its real duration and truncates to whole seconds (10.9 seconds counts as 10); anything under 1 second counts as 1. This matches the upstream’s own accounting.
What happens if it cannot be read — if the header cannot be parsed (unusual format, unreachable URL, and so on) the request is rejected with 400 input_video_duration_unknown and nothing is charged — it is not estimated at a cap and billed. Use a directly downloadable URL in a common container format (mp4 / webm).
Estimating before you submit — POST https://api.qingbo.ai/waveapi/quote with video_urls returns the exact price for that call, including the input seconds it probed.
This is a public utility endpoint and needs no API key. It replies in the standard envelope {code, message, data}, where data carries:
When a request is rejected,
errorKey carries the specific reason (such as input_video_duration_unknown) and message holds text you can show to the user directly.
How the reference video then enters the price (counted as seconds, only switching tier, or replacing the clip price) depends on the vendor — see the billing section on each model page.
Submit Response Example
task_id, poll GET /v1/tasks/{task_id} until status = completed, then read the video URL.
Mode Quick Reference
Not every model supports every mode — check each vendor’s doc for the actual supported
action list and field range. The backend validates whether request fields fall within the vendor’s declared capabilities and returns an error if they do not.Field Naming Conventions
- Media references are always plural — always
image_urls/video_urls/audio_urls; even a single video uses a one-element array["one.mp4"] - Frame ratio is unified as
aspect_ratio— vendor-internalsize/ratioetc. are implementation details and not exposed - Resolution is unified as
resolution— vendor-internalmode/qualityetc. are implementation details - First/last-frame fields carry the
_imagesuffix —first_frame_image/last_frame_image(emphasizing the image resource)
Related Docs
- Task System — task state machine / polling cadence / webhook
- Request and Response Format — common error codes / headers / rate limits
- Authentication — API key application and usage
- Model List — full model list query endpoint