Skip to main content
Model launch
7 new models · text / image / video
  • Text · Anthropic: claude-haiku-5-5 — the smallest model of Claude 5.5, 1M-token context, adaptive thinking on by default, image input, function calling and the native /v1/messages endpoint; a request with more than 100K input tokens is billed entirely at the long-context tier.
  • Image · Black Forest Labs: flux-3-image — FLUX 3 Image: text-to-image, single-image editing and up to 10 reference images, five output tiers up to 4K, and bounding boxes in the prompt to steer layout and local edits; billed per image, reference images are free.
  • Image · Google: gemini-nano-banana-2.1 and gemini-nano-banana-2.1-rev — Nano Banana 2.1 with 1K / 2K / 4K output and up to 14 reference images. The official line settles on actual token usage and returns up to 4 images per request; the economy line has a fixed per-image price and returns one image.
  • Image · Microsoft: mai-image-2.6 and mai-image-2.6-flash — MAI-Image-2.6 and its faster version: text-to-image, single-image editing and composition from up to 5 references, 1K / 2K, exact pixel sizes and optional web grounding; billed on actual token usage.
  • Video · Vidu: vidu-q4-preview — Vidu Q4 Preview: start-frame video, or video from 1–15 reference images plus up to 3 reference audio clips, 3–16 seconds, 540p to 4K, with audio by default.
  • Docs: Text models overview · FLUX 3 Image · Nano Banana 2.1 · Nano Banana 2.1 Economy · MAI-Image-2.6 · Vidu Q4 Preview
Model launch
Text models · 2 new models (OpenAI / Anthropic)
  • OpenAI: gpt-6.1-sol — GPT-6.1 Sol, about 1M-token context, reasons before answering, image input, JSON Schema structured output and function calling (function calling goes through /v1/responses); reasoning_effort accepts low / medium / high / xhigh; prompts over 272K tokens are billed at the long-context tier.
  • Anthropic: claude-sonnet-5-5 — the Sonnet model of Claude 5.5, 1M-token context, adaptive thinking on by default, image input, function calling, web search and the native /v1/messages endpoint, priced the same as claude-sonnet-5.
  • Both served via /v1/chat/completions and /v1/responses with streaming.
  • Docs: Text models overview
Capability update
Asset library launched · Seedance asset references
  • New /v1/assets endpoints: upload image, video and audio links as assets; once approved, reference them in generation requests as asset://<asset ID>.
  • Supported models: seedance-2.0 · seedance-2.0-fast · seedance-2.0-mini · seedance-2.5. Real-person portraits used as references must be uploaded as assets first.
  • Asset endpoints are free; up to 500 assets per account.
  • Docs: Asset library · Seedance series
Model launch
MiniMax H3 Max launched · Video
  • Model ID: minimax-h3-max
  • The high-quality tier of MiniMax H3: 480P / 768P / 1080P, with reference images, reference video and reference audio.
  • Docs: MiniMax H3
Capability update
Video models: Veo / Kling 4K, SkyReels / PixVerse extension, FLUX 3 Video finalize, Grok video editing
  • Veo 3.1: 4K added to veo3.1-fast / veo3.1-quality / veo3.1-lite / veo3.1-fast-rev / veo3.1-quality-rev; economy extensions accept resolution (a 4K extension needs a 4K parent); enable_gif cannot be combined with 1080p / 4K.
  • Kling: 4K on kling-v3 / kling-v3-omni; kling-v3-omni takes up to 7 reference images, cited in the prompt as <<<image_N>>>.
  • SkyReels V4: 1080p with a reference video; new extend action.
  • PixVerse V6: new extend action using ref_task_id; default resolution is now 540p.
  • FLUX 3 Video: flux-3-video adds finalize to turn a flux-3-video-draft draft into the final video.
  • Grok Imagine Video: grok-imagine-video adds edit; grok-imagine-video-1.5 accepts multiple reference images (an explicit aspect ratio is required).
  • Seedance: seedance-2.5 accepts duration: 0 for a model-chosen length; on seedance-1.5-pro and the three 2.0 tiers, image_with_roles cannot be combined with image_urls.
  • Wan: wan3.0-video / wan3.0-video-prime support a model-chosen length; wan2.7-r2v reference clips 1–30 s each; wan2.5-preview always generates audio.
  • Hailuo: minimax-hailuo-2.3 / -2.3-fast no longer accept seed.
  • Vidu: vidu-q3 / vidu-q3-mix take up to 7 reference images; all four models default to 720p.
  • Docs: Veo series · Kling series · SkyReels V4 series · PixVerse V6 · FLUX 3 Video · Grok Imagine Video series · Seedance series · Wan series · Hailuo series · Vidu Q3 series
Model delisted
Sora 2 series delisted · Video
  • Model IDs: sora-2 · sora-2-pro
  • OpenAI shut down every Sora 2 model and the Videos API on 2026-09-24, with no official replacement. These models are delisted effective today and calls will return an error.
  • For video generation, use Veo 3.1, Kling, Seedance or another video model.
  • Docs: Video models overview
Model launch
Seedream 5.0 Flash launched · Image
  • Model ID: seedream-5.0-flash
  • The fast tier of ByteDance’s Seedream 5.0 line: text-to-image and editing with up to 10 reference images, 1k / 1.5k / 2k, automatic framing plus 10 common aspect ratios, JPEG / PNG output, and transparent-background editing; one image per request.
  • Fixed price per image, the same at every resolution; reference images are free.
  • Docs: Seedream series
Capability update
Video models: SkyReels default resolution, Wan 2.7 video-edit billing, Gemini Omni duration, Grok Imagine Video reference images
  • skyreels-v4-fast / skyreels-v4-std: resolution now defaults to 1080p (the resolution actually rendered); with a reference video (video_urls) you must send 480p or 720p explicitly, otherwise the request returns 400.
  • wan2.7-videoedit: the source video’s seconds are now charged in addition to output seconds.
  • gemini-omni-1.1-flash / gemini-omni-flash-preview: duration must be omitted or 0; the model picks the length (3–10 s).
  • grok-imagine-video: reference images (image_urls) are now charged per image.
  • grok-imagine-1.0-video-rev is retired and calls will return an error; use grok-imagine-video-1.5-rev instead.
  • Docs: SkyReels · Wan series · Gemini Omni Flash · Grok Imagine Video
Capability update
Image models: GPT Image 2.5 moves to the official model ID and gains a reverse line; several models now return multiple images
  • gpt-image-2.5 is renamed gpt-image-2.5-flare to match OpenAI’s official model ID; parameters and prices are unchanged. The old ID is disabled and now returns model-not-found — change model to gpt-image-2.5-flare.
  • Inpainting on GPT Image 2.5: gpt-image-2.5-flare and gpt-image-2.5-sunburst accept mask_url (a PNG with an alpha channel; image_urls required).
  • New GPT Image 2.5 reverse line: gpt-image-2.5-flare-rev · gpt-image-2.5-sunburst-rev — text-to-image and reference editing, 1k / 2k / 4k, one image per request, fixed per-image price by resolution tier.
  • Multiple images per request: grok-imagine-image-2.0 returns 1–10 images, grok-imagine-image-2.0-rev 1–12, qwen-image-3.0 / qwen-image-3.0-pro 1–6, seedream-5.0-lite up to 15, and seedream-4.0 / seedream-4.5 up to 15 reference plus output images when editing. Multi-image requests are billed by the number of images actually returned.
  • z-image-turbo no longer accepts prompt_extend; sending it returns 400 and is not billed.
  • Docs: GPT Image series · GPT Image 2.5 Reverse · Grok Imagine · Qwen Image · Seedream
Capability update
45 text models support the Responses API
  • Text models from OpenAI (the gpt-4.1 / gpt-5 / gpt-5.x / gpt-6 lines, o1 / o3 / o3-mini / o4-mini), Qwen, DeepSeek and xAI can now also be called through POST /v1/responses, with streaming and function calling, so clients such as Codex can use them directly. deepseek-v3.1-terminus is not supported yet.
  • Function calling is now available on gpt-6-astra and grok-4.6 through /v1/responses.
  • Billing is the same as on the Chat endpoint.
  • Docs: Text models overview
Model launch
Text models · 10 new models (OpenAI / Anthropic / xAI / Qwen / Z.ai)
  • OpenAI: gpt-6-sol · gpt-6-luna — the mid and low-cost tiers of GPT-6, about 1M-token context, image input, JSON Schema structured output and function calling (function calling requires reasoning_effort: "none"); prompts over 272K tokens are billed at the long-context tier.
  • Anthropic: claude-opus-5-5 — the first Claude 5.5 model, 1M-token context, always thinks, supports function calling and the native /v1/messages endpoint, priced below claude-opus-5.
  • xAI: grok-4.7 — 500K-token context, image input and function calling, always reasons; prompts of 200K tokens or more are billed at the long-context tier.
  • Qwen: qwen3.8-max-0902 (upgraded snapshot of qwen3.8-max, released 2 September) · qwen3.8-2.4t-a95b (open-weight flagship, text input only) · qwen3.8-flash · qwen3.8-27b · qwen3.7-flash (long-context tier above 32K input).
  • Z.ai: glm-5.3 — always reasons, supports function calling.
  • All served via /v1/chat/completions with streaming.
  • Docs: Text models overview
Model launch
DeepSeek V4.1 Flash is available
  • Model ID: deepseek-v4.1-flash
  • The next generation of DeepSeek V4 Flash: a 1M-token context, up to about 384K output tokens per response, and thinking on by default.
  • Served on POST /v1/chat/completions with streaming, JSON Schema structured output and function calling; tokens that hit the automatic cache settle at the cache-read rate.
  • Time-of-day pricing: a request settles at the rates of the window it was submitted in. The full schedule is price_config.text_schedule on GET /v1/models.
  • Docs: Text Models Overview
Model launch
Wan 3.0 Video Prime is available
  • Model ID: wan3.0-video-prime
  • The accelerated edition of wan3.0-video, with faster generation. Actions, parameters, value ranges and constraints are exactly the same as wan3.0-video: text, first/last frames, reference images, videos and audio, plus document (file_url) and web-page (link_url) references; 480p / 720p / 1080p, 2–30 seconds. Switching only changes model.
  • Billed by resolution tier × output seconds at a higher rate than wan3.0-video; with no resolution, the task settles at the 480p tier.
  • Docs: Wan Series
Model launch
GPT Image 2.5 is available, adding the xhigh and max quality levels
  • Model IDs: gpt-image-2.5 (the Flare line, faster output) · gpt-image-2.5-sunburst (the Sunburst line, editing precision first)
  • Both models share the same actions, parameters, value ranges and rates: text-to-image and reference editing (up to 16 reference images), 15 aspect ratios, 1k / 2k / 4k, one image per request. Only the generated result differs.
  • Quality widens from three levels to five: low / medium / high / xhigh / max. xhigh and max exist only on 2.5; sending them to gpt-image-2 returns 400 and is not downgraded silently.
  • Levels sharing a name are not interchangeable across generations: medium and high on 2.5 are about a quarter of the output tokens of the same-named levels on GPT Image 2, and max on 2.5 is what corresponds to high on GPT Image 2. Re-estimate cost against the new page when migrating from gpt-image-2.
  • Billed from actual token usage across text input, cached text input, image input, cached image input and image output. This line does not accept mask_url or seed; masked inpainting stays on gpt-image-2.
  • Docs: GPT Image 2.5 · GPT Image Series
Capability update
Multi-image generation opened on Wan 2.7 Image
  • Model IDs: wan2.7-image · wan2.7-image-pro
  • n is no longer fixed at 1: 1–4 images per request in standard mode, 1–12 with enable_sequential (one thematically coherent set per request). wan2.7-image-pro tops out at 2k in sequential mode.
  • Billing is unchanged: a fixed price per image, n images cost n times the unit price, and the hold at submission is sized by n.
  • Docs: Wan Image Series
Model launch
Nano Banana Official is live · one more economy tier
  • Model IDs: gemini-3-pro-image (flagship) · gemini-3.1-flash-image (fast) · gemini-3.1-flash-lite-image (lightest) · gemini-2.5-flash-image-preview-rev (previous-generation economy tier)
  • The three official tiers are served on POST /v1/tasks with generate and edit (up to 14 reference images), one image per request. The flagship offers 1k / 2k / 4k; the fast tier offers 0.5k / 1k / 2k / 4k plus the extreme ratios 1:4, 4:1, 1:8 and 8:1; the lightest always outputs 1K and rejects resolution.
  • The official line is billed from actual token usage: a hold is placed for the estimated usage at submission, the task settles from actual usage, and the difference is released. The dimensions are text input, image input, text output and image output, with reference images counting as image input.
  • google_search / google_image_search are accepted only by gemini-3.1-flash-image and gemini-3.1-flash-image-rev.
  • The economy line adds gemini-2.5-flash-image-preview-rev: 11 aspect ratios, fixed 1K, still a fixed price per successful image.
  • Docs: Gemini Image Series · Nano Banana Official · Nano Banana Economy
Model launch
GPT Image 1 and GPT Image 1.5 are available again
  • Model IDs: gpt-image-1 · gpt-image-1.5
  • Both share one contract: text-to-image and reference editing, masked inpainting, transparent backgrounds, png / jpeg output, aspect ratios 1:1 / 2:3 / 3:2, four quality tiers, up to 15 reference images, one image per request. Neither accepts resolution or seed.
  • Billed from actual token usage; gpt-image-1.5 bills text output on top.
  • Docs: GPT Image Series · GPT Image 1 and 1.5 Official
Model launch
Grok Imagine · four image tiers and the 1.5 reverse video tier
  • Model IDs: grok-imagine-image-2.0 · grok-imagine-image-2.0-rev · grok-imagine-image-quality · grok-imagine-1.5-rev · grok-imagine-1.5-edit-rev · grok-imagine-video-1.5-rev
  • Image: grok-imagine-image-2.0 official and -rev reverse tiers; grok-imagine-image-quality is xAI’s official quality tier with 14 aspect ratios, 1k / 2k tiers and editing with 1–3 reference images, each reference image billed separately; grok-imagine-1.5-rev (text-to-image) and grok-imagine-1.5-edit-rev (generate / edit) return up to 10 images per request. All are priced per image.
  • Video: grok-imagine-video-1.5-rev does text- and image-to-video at 480p / 720p, 6–15 seconds, up to 7 reference images, billed by resolution tier × seconds.
  • Docs: Grok Imagine Series · Grok Imagine Video Series
Model launch
Gemini Omni Flash is live · video generation with a model-chosen length
  • Model IDs: gemini-omni-1.1-flash · gemini-omni-flash-preview
  • Both run generate on POST /v1/tasks and build video from text, reference images or one reference clip. Neither takes duration — the model decides the length from the content. Aspect ratios 16:9 / 9:16, one video per request.
  • gemini-omni-1.1-flash: 360p / 720p / 1080p / 4k, 10 reference images in total (frames included), first_frame_image / last_frame_image, image_with_roles and metadata, reference video up to 10 s. gemini-omni-flash-preview: 720p only, up to 16 reference images, reference video up to 24 s.
  • Both accept ref_task_id: pass the previous task’s task_id to extend or edit that result. It is mutually exclusive with video_urls.
  • Billed on resolution tier × output seconds. A hold is placed at the maximum length on submission and settled on the actual result, with the difference released.
  • Docs: Omni Video Series · Gemini Omni Flash · Omni-Flash-Ext
Model launch
Four video models added · Wan 3.0, HappyHorse 1.1 and both FLUX 3 Video tiers
  • Model IDs: wan3.0-video · happyhorse-1.1 · flux-3-video · flux-3-video-draft
  • wan3.0-video: the Wan multimodal video entry point — one model ID for text, first/last frames, reference images, videos and audio, plus document (file_url) and web-page (link_url) references; 480p / 720p / 1080p, 2–30 seconds; generation_type selects the frame or reference input family.
  • happyhorse-1.1: one entry point for text-to-video, first-frame image-to-video and reference-image video; 720p / 1080p, any integer 3–15 seconds, five aspect ratios.
  • flux-3-video / flux-3-video-draft: 5–20 second clips with synchronized audio from text, 1–10 ordered keyframes, or a continuation video, across eight aspect ratios including auto; the standard tier takes hd and fhd, the Draft tier is hd only at lower quality and lower cost for previews.
  • The edit action on happyhorse-1.0 now has explicit duration semantics: duration: 0 means “use the source video length”, and billed seconds come from the source video.
  • Docs: Wan Series · HappyHorse Series · FLUX 3 Video
Model launch
Three video models added · Hailuo 02, Vidu Q3 Pro and Turbo
  • Model IDs: minimax-hailuo-02 · vidu-q3-pro · vidu-q3-turbo
  • minimax-hailuo-02: MiniMax Hailuo 02 — text-to-video, first frame and first-and-last frame at 512p / 768p / 1080p, 5 or 10 seconds. 512p needs a first frame and 1080p is 5 seconds only.
  • vidu-q3-pro / vidu-q3-turbo: the quality and speed tiers of Vidu Q3, with identical capabilities — text, start frame or first-and-last frame at 540p / 720p / 1080p, 1–16 seconds, native audio on by default (generate_audio). With input images the aspect ratio follows the image.
  • All three use the generate action on POST /v1/tasks and bill by resolution tier × seconds.
  • Docs: Hailuo Series · Vidu Q3 Series
Model launch
Flow Music is live · music generation
  • Model ID: flow-music
  • Text-to-music on the Lyria engine: generate produces a full song or an instrumental (with lyrics, title, BPM, length and seed), and lyrics writes lyric text on its own, both through POST /v1/tasks. Billed per call, with different prices for the two actions.
  • Suno adds the download action in the same batch, exporting a source track as an audio file.
  • Docs: Flow Music · Suno Music Generation
Capability update
Video edit and extend on Seedance 2.5
  • Model ID: seedance-2.5
  • Two new actions: edit rewrites a source video from the prompt, with duration: 0 keeping the source length; extend continues a source video. Both take the source in video_urls and require the adaptive aspect ratio.
  • The Seedance, Sora, PixVerse and SkyReels pages were re-checked against the catalog; seedance-1.0-pro and seedance-1.0-pro-fast durations are corrected to any integer from 2 to 12 seconds.
  • Docs: Seedance Series · Sora 2 Series · PixVerse V6 · SkyReels V4 Series
Capability update
Kling and Veo actions and parameters aligned
  • New actions: kling-v3-omni and kling-video-o1 support action: "edit", editing a 3–10 second source video and keeping its duration; veo3.1-fast-rev and veo3.1-quality-rev support action: "extend", continuing their own earlier video via ref_task_id.
  • New parameters: Kling multi-shot storyboards multi_shot / shot_type / multi_prompt and reference subjects element_list (kling-v3, kling-v3-omni), role-tagged images image_with_roles (kling-v3-omni, kling-video-o1); on Veo, person_generation / resize_mode / seed / negative_prompt on the official tiers and generation_type / enable_gif / raw on the economy tiers.
  • Duration: kling-3.0-turbo, kling-v3 and kling-v3-omni take any integer from 3 to 15 seconds; veo3.1-fast and veo3.1-quality take 4, 6 or 8 seconds, with 1080p at 8 seconds only.
  • Docs: Kling Series · Veo Series
Capability update
Midjourney chained actions completed
  • Model ID: midjourney
  • The action table now covers 12 chained actions: upscale, variation, high-variation, low-variation, reroll, zoom, pan, inpaint, modal, edit, remix-strong and remix-subtle, all referencing the previous step through ref_task_id and selecting a grid image with index.
  • The page also documents the full imagine parameter set (stylize, chaos, weird, tile, iw, cref / sref / dref with their weights, raw, draft, hd, stop, extra) and the speed / version values (v8.2 added); billing is by action and speed tier.
  • Docs: Midjourney
Model retirement
5 media models retired
  • Model IDs: seedance-2.0-face · seedance-2.0-fast-face · minimax-h3-max · grok-imagine-1.0-rev · grok-imagine-1.0-edit-rev
  • seedance-2.0-face / seedance-2.0-fast-face: seedance-2.0 and seedance-2.0-fast now accept face references themselves, so the separate face tiers are gone. Switch to the model IDs without the -face suffix.
  • minimax-h3-max: no longer offered upstream. The MiniMax H3 line continues with minimax-h3 and minimax-h3-regeneration.
  • grok-imagine-1.0-rev / grok-imagine-1.0-edit-rev: no longer offered upstream. Use grok-imagine-1.5-rev / grok-imagine-1.5-edit-rev instead.
  • These models are retired as of today; calls to them return an error.
  • Docs: Seedance Series · MiniMax H3 · Grok Imagine Series
Model launch
Nano Banana Economy is live · three Gemini image tiers
  • Model IDs: gemini-3-pro-image-rev (flagship) · gemini-3.1-flash-image-rev (fast) · gemini-3.1-flash-lite-image-rev (lightest)
  • All three share one parameter set: served on POST /v1/tasks with generate and edit (up to 14 reference images), one image per request, and eleven shared aspect ratios including auto.
  • Tier differences: the flagship charges the same at 1k and 2k; the fast tier costs more at 2k than 1k and adds the extreme ratios 1:4, 4:1, 1:8 and 8:1; the lightest is the cheapest, always outputs 1K and rejects resolution. Rates are on GET /v1/models.
  • Billed at a fixed price per successful image. Resolution is the only billing dimension — reference images and aspect ratio add nothing, so cost is known before submission.
  • The economy routes have no line-specific parameters: mask_url, quality, seed, google_search / google_image_search and official_fallback are all rejected before quota is reserved. Use gpt-image-2 for masked inpainting.
  • Docs: Nano Banana Economy
Model launch
GPT Image 2 is live · official and reverse lines
  • Model IDs: gpt-image-2 (official) · gpt-image-2-rev (reverse)
  • Both are served on POST /v1/tasks with generate and edit, fifteen aspect ratios, 1k / 2k / 4k resolutions, one image per request, and no seed support.
  • The official line adds masked inpainting (mask_url), quality tiers (low / medium / high) and transparent-background and output-format control (background / output_format / output_compression / moderation), with up to 16 reference images. It settles from actual token usage: the reservation is estimated from aspect ratio × resolution × quality, and the difference is refunded on completion.
  • The reverse line accepts none of those line-specific parameters and takes up to 15 reference images. It bills a fixed price per image by output resolution, with different rates at 1k, 2k and 4k, so the cost of a call is known before you send it.
  • The two lines never switch automatically — pin the model ID in your configuration as a product choice.
  • Docs: GPT Image Series · GPT Image 2 Official · GPT Image 2 Reverse
Model launch
gpt-4o-mini-tts is live · Text-to-speech
  • Model ID: gpt-4o-mini-tts
  • Replaces the delisted tts-1 / tts-1-hd. Thirteen voices (alloy / echo / fable / onyx / nova / shimmer / coral / verse / ballad / ash / sage / marin / cedar), plus a new instructions parameter for setting tone, pace and emotion in plain language.
  • Billed by input text characters at $13.5 per million characters, with no output charge. Served via /v1/audio/speech.
  • Docs: Text-to-speech
Capability update
Reference-video input is now available across the video models
  • Reference video supported: kling-v3-omni · kling-video-o1 · skyreels-v4-fast · skyreels-v4-std · gemini-omni-1.1-flash-ext
  • Relisted: kling-v2-6-motion-control · kling-v3-motion-control
  • seedance-2.0-fast now accepts reference videos, matching seedance-2.0 and seedance-2.0-mini.
  • Input video duration is billed on the real file duration (truncated to whole seconds); when it cannot be parsed the request is rejected and not billed. How each vendor folds the reference video into the price is on the model pages; the shared rules are in Input video duration and billing.
  • Docs: Kling series · SkyReels V4 · Omni-Flash-Ext · Seedance series
Model delisted
Seven models delisted
  • Model IDs: tts-1 · tts-1-hd · sora-2-preview · deepseek-v3-0324 · deepseek-r1-250528 · kimi-k2-instruct · gpt-4o-2024-08-06
  • Upstream no longer serves these models; they are delisted effective today and calls will return an error.
  • tts-1 / tts-1-hd: text-to-speech will return as gpt-4o-mini-tts.
  • gpt-4o-2024-08-06: gpt-4o-2024-05-13 / gpt-4o-2024-11-20 remain available.
Model delisted
Imagen 4.0 · Image
  • Model ID: imagen-4.0
  • Upstream no longer serves this model; it is delisted effective today and calls will return an error.
  • For Google image generation, use the Nano Banana family instead: gemini-3-pro-image-preview (Nano Banana Pro) / gemini-3.1-flash-image-preview (Nano Banana 2), plus their -rev standard tiers.
  • Docs: Gemini image models
Model launch
Nano Banana 2 Lite standard tier · Image
  • Model ID: gemini-3.1-flash-lite-image-rev
  • The fastest, cheapest image model in the Gemini 3.1 family — about 4 seconds per image, 1K only, up to 14 reference images, no search grounding or inpainting. Currently the lowest-priced Gemini image tier on the platform.
  • Docs: Gemini image models
Model launch
Text models · 15 new models (Anthropic / Google / Qwen / MiniMax / MoonshotAI / Z.ai)
  • Anthropic: claude-fable-5 — Mythos-class flagship, 1M-token context, no long-context surcharge.
  • Google: gemini-3.7-flash · gemini-3.6-flash · gemini-3.5-flash-lite.
  • Qwen (first Qwen text models on the platform): qwen3.8-max (released today) · qwen3.7-max · qwen3.7-plus · qwen3.6-plus · qwen3.6-flash, all with 1M-token context; Plus / Flash bill a long-context tier above 256K input.
  • MiniMax: minimax-m3 (multimodal, 1M context, 2× above 512K input) · minimax-m2.7.
  • MoonshotAI: kimi-k2.6 · kimi-k2.7-code · kimi-k2.7-code-highspeed.
  • Z.ai: glm-5.2 (1M context).
  • All served via /v1/chat/completions; model IDs are bare names with no vendor prefix.
  • Docs: Chat completions
Model launch
Grok text models · four models
  • Model IDs: grok-4.6 · grok-4.5 · grok-4.3 · grok-build-0.1
  • xAI’s current-generation text line is now available: 4.6 is the flagship (code and everything else, 500K context); 4.5 is the reasoning tier; 4.3 is the fast general-purpose tier (1M context); Build 0.1 targets coding agents. All support image input, function calling and structured outputs via /v1/chat/completions.
  • Requests with more than 200K input tokens are billed at 2× the standard rate, matching xAI’s long-context tier.
  • In the same batch, kimi-k3 output pricing was re-aligned (lowered) to Moonshot’s official list price.
  • Docs: Chat completions
Model launch
GPT-4o Transcribe (two tiers) · Speech to text
  • Model IDs: gpt-4o-transcribe · gpt-4o-mini-transcribe
  • OpenAI’s current-generation transcription models with a markedly lower word-error rate than Whisper; the Mini tier is half price. Billed per minute of audio via /v1/audio/transcriptions.
  • Docs: Speech transcription
Model launch
MiniMax H3 Max / Regeneration · Video
  • Model IDs: minimax-h3-max (480p / 768p value tier) · minimax-h3-regeneration (re-renders a 768p clip at native 2K, video_urls required)
  • minimax-h3 pricing was aligned to MiniMax’s official rate card in the same batch.
  • Docs: MiniMax H3
Model launch
Video · Seedance 1.0 Pro (two tiers) + Wan 2.6 image-to-video (two tiers)
  • Model IDs: seedance-1.0-pro · seedance-1.0-pro-fast · wan2.6-i2v · wan2.6-i2v-flash
  • Seedance 1.0 Pro is the stable 1.x flagship, priced below 1.5 / 2.x; the Fast tier is about 3× faster and the cheapest Seedance video option. 480p / 720p / 1080p, 5 or 10 sec.
  • Wan 2.6 image-to-video now has a standard and a Flash tier (about half price), 720p / 1080p, 5 / 10 / 15 sec.
  • Billing correction: across the Seedance family, image-to-video costs the same as text-to-video (the docs previously described two tiers); reference video is billed at 2× the unit price as an interim rule. Wan 2.7 video editing and reference-to-video use the same 2× interim rule.
  • Docs: Seedance · Wan series
Model launch
Seedream 4.0 / 4.5 and Imagen 4 · Image
  • Model IDs: seedream-4.0 · seedream-4.5 · imagen-4.0
  • Both Seedream 4.x tiers are now listed: 4.0 is the lowest-priced Seedream, 4.5 has better text accuracy and multi-reference fusion; both support 1K / 2K / 4K, up to 10 reference images and 15 images per request.
  • Imagen 4 (standard tier) is now open — the docs previously flagged it as under maintenance. Photorealistic text-to-image, 1K / 2K, five aspect ratios, text-to-image only.
  • Docs: Seedream series · Imagen 4.0
Model launch
FLUX.2 Max · Image
  • Model ID: flux-2-max
  • The top quality tier of the FLUX.2 family, for final renders where fidelity outweighs cost. Up to 8 reference images, 1K / 2K, billed by output megapixels.
  • Docs: FLUX series
Model launch
Grok Imagine official line · four models
  • Model IDs: grok-imagine-image (standard) · grok-imagine-image-quality (quality) · grok-imagine-video (1.0) · grok-imagine-video-1.5
  • Until now only the Grok Imagine economy line (-rev) was listed; the xAI official line is now available too. The standard image tier costs the same at 1K and 2K and is the cheapest official Grok image model; Video 1.5 adds 1080p and audio-driven generation.
  • Correction: the Grok video duration range is 6-15 seconds; the docs previously said 6-30.
  • Docs: Grok Imagine Image · Grok Imagine Video
Model launch
Qwen Image 3.0 · Image
  • Model IDs: qwen-image-3.0 (standard) · qwen-image-3.0-pro (dense layouts)
  • Markedly more reliable text rendering than gen-2. 1K / 2K, seven aspect ratios, up to 3 reference images, up to 6 images per request.
  • The Pro tier prices 1K and 2K separately; the standard tier charges the same for both.
GPT Image 1.5 · Image
  • Model ID: gpt-image-1.5
  • Faster than GPT Image 1, with markedly better instruction following and editing precision. 1:1 / 2:3 / 3:2, four quality tiers.
  • Omitting quality bills at the high rate — see GPT-Image series.
Speech · three models
  • Model IDs: tts-1 · tts-1-hd · whisper-1
  • TTS is billed per character of input text; Whisper is billed per minute of audio.
Breaking change
Seedance / Seedream · doubao- prefix dropped from model IDsFor consistency with the rest of the catalog — most model IDs here carry no vendor prefix, and vendor is shown separately on the model card — these 7 model IDs have been renamed:
  • The old IDs no longer resolve. Please switch to the new IDs.
  • Capabilities, parameters and pricing are unchanged — this is a rename only.
  • Docs: Seedance Video API · Seedream Image API
Model launch
Claude Sonnet 5 · Chat
  • Model ID: claude-sonnet-5
GPT-5.6 series · Chat
  • Model IDs: gpt-5.6-luna (light) · gpt-5.6-terra (standard) · gpt-5.6-sol (flagship)
  • Three tiers with increasing capability and price — pick the one that fits your workload.
Capability update
Omni Flash Ext · Now billed by resolution × duration
  • omni-flash-ext moves from a single flat rate to resolution × duration tiers: resolution 720p / 1080p / 4k, duration 4 / 6 / 8 / 10 seconds (only these four values are accepted).
  • When passing a reference video (video_urls), do not pass duration — the two are mutually exclusive.
  • Current unit prices are shown in the Models overview and the console “Model Market”.
Breaking change
Seedance 2.0 series · image_urls semantics changeEffective 2026-07-09 (upstream), for seedance-2.0, seedance-2.0-fast and seedance-2.0-mini:
  • Before — the 1st image in image_urls acted as the first frame, the 2nd as the last frame, the rest as references.
  • After — all images in image_urls are treated as reference images only; they no longer carry first/last-frame meaning.
To keep constraining first / last frames, use image_with_roles (unaffected by this change). seedance-1.5-pro and the face variants are unaffected.
Capability update
Seedance 2.0 · 4K added
  • seedance-2.0 now supports 4K — 480p / 720p / 1080p / 4K; pricing synced to official rates.
  • Docs: Seedance video API
Model launch
HappyHorse 1.1 · Video generation
  • Model ID: happyhorse-1.1
  • Modes: text-to-video, image-to-video (pass image_urls as the first frame)
  • Resolution: 720P / 1080P · Duration: 3–15 s · Aspect ratios: 16:9 / 9:16 / 1:1
  • Docs: HappyHorse video API
Model launch
Kling 3.0 Turbo · Video generation
  • Model ID: kling-3.0-turbo
  • Modes: text-to-video, image-to-video (pass image_urls as the first frame)
  • Resolution: 720P / 1080P · Duration: 5 / 10 s · Aspect ratios: 16:9 / 9:16 / 1:1
  • Docs: Kling video API
Grok Imagine 1.5 · Image / edit / video
  • Image grok-imagine-1.5-rev — text-to-image, aspect ratios 1:1 / 16:9 / 9:16 / 3:2 / 2:3, batch n
  • Edit grok-imagine-1.5-edit-rev — generate / edit, pass image_urls as references
  • Video grok-imagine-1.5-video-rev — text/image-to-video, resolution 480p / 720p, duration 6–30 s, ratios 16:9 / 9:16 / 1:1
  • Docs: Grok Imagine image API · Grok Imagine video API
Model launch
Seedance 2.0 Mini · Video generation
  • Model ID: seedance-2.0-mini
  • Modes: text-to-video, image-to-video
  • Resolution: 480p / 720p · Duration: 4–15 s · Aspect ratios: 16:9 / 9:16 / 1:1
  • Docs: Seedance video API
Model launch
Claude Opus 4.8 · Text generation
  • Model ID: claude-opus-4-8
  • Access: OpenAI-compatible Chat Completions, streaming supported
  • Docs: Text generation API
Model launch
Gemini 3.5 Flash · Text generation
  • Model ID: gemini-3.5-flash · OpenAI-compatible Chat Completions
  • Docs: Text generation API
Omni-Flash-Ext · Video generation
  • Model ID: omni-flash-ext
  • Modes: text-to-video, video-reference generation (pass video_urls)
  • Resolution: 720p / 1080p / 4K · Duration: 4 / 6 / 8 / 10 s · Aspect ratios: 16:9 / 9:16 / 1:1
  • Docs: Omni video API
Model launch
Midjourney · Image generation
  • Model ID: midjourney
  • Modes: text-to-image (imagine), blend, describe, edits; chained upscale / variation / zoom / pan (pass ref_task_id)
  • Aspect ratios: 1:1 / 16:9 / 9:16 / 2:3 / 3:2 / 4:3 / 3:4 / 21:9 / 9:21 · Speed tiers: relax / fast / turbo
  • Docs: Midjourney image API
Model launch
Gemini Flash Latest · Text generation
  • Model ID: gemini-flash-latest · rolling alias pointing to the latest Gemini Flash, OpenAI-compatible Chat Completions
  • Docs: Text generation API
Model launch
GPT-5.5 · Text generation
  • Model ID: gpt-5.5 · next-gen flagship text model, OpenAI-compatible Chat Completions
  • Docs: Text generation API
DeepSeek V4 · Text generation
  • deepseek-v4-pro — for deep reasoning, complex multi-turn, and long-context tasks
  • deepseek-v4-flash — lightweight and ultra-fast, for high-concurrency, low-latency interactive use
  • Docs: Text generation API
Model launch
GPT Image 2 series
  • Model ID: gpt-image-2-rev · gpt-image-2
  • Resolution 1K / 2K / 4K · Aspect ratios 15 ratios · generate / edit
  • Docs: GPT Image API
Model launch
Claude Opus 4.7
Model launch
Vidu Q3 series
  • Model ID: vidu-q3 · vidu-q3-mix
  • Resolution 540p / 720p / 1080p · Duration 1–16 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Vidu video API
Model launch
GLM-5.1
Model launch
Wan 2.7 · Video
  • Model ID: wan2.7 · wan2.7-r2v · wan2.7-videoedit
  • Resolution 720P / 1080P · Duration 2–15 s · Aspect ratios 16:9 / 9:16 / 1:1 / 4:3 / 3:4
  • Docs: Wan video API
Wan 2.7 · Image
  • Model ID: wan2.7-image · wan2.7-image-pro
  • Resolution 1K / 2K / 4K · Aspect ratios 7 ratios · generate / edit
  • Docs: Wan image API
Model launch
PixVerse V6
  • Model ID: pixverse-v6
  • Resolution 360p / 540p / 720p / 1080p · Duration 1–15 s · Aspect ratios 8 ratios
  • Docs: PixVerse video API
Model launch
GPT-5.4 mini / nano
Model launch
GPT-5.4
Model launch
Gemini 3.1 Flash Lite Preview
Model launch
Gemini 3.1 Flash Image Preview
  • Model ID: gemini-3.1-flash-image-rev · gemini-3.1-flash-image
  • Resolution 0.5K / 1K / 2K / 4K · Aspect ratios 11 ratios · generate / edit
  • Docs: Gemini image API
Model launch
Gemini 3.1 Pro Preview (custom tools)SkyReels V4
  • Model ID: skyreels-v4-std · skyreels-v4-fast
  • Resolution 480p / 720p / 1080p · Duration 3–15 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: SkyReels video API
Model launch
GPT-5.3 Codex
Model launch
Gemini 3.1 Pro Preview
Model launch
Claude Sonnet 4.6
Model launch
Seedream 5.0 Lite
  • Model ID: seedream-5.0-lite
  • Resolution 2K / 3K / 4K · Aspect ratios 9 ratios · generate / edit
  • Docs: Seedream image API
Model launch
MiniMax M2.5Seedance 2.0 series
  • Model ID: seedance-2.0 · seedance-2.0-fast · seedance-2.0-face · seedance-2.0-fast-face
  • Resolution 480p / 720p / 1080p · Duration 4–15 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Seedance video API
Model launch
GLM-5
Model launch
Qwen Image 2.0
  • Model ID: qwen-image-2.0 · qwen-image-2.0-pro
  • Resolution 1K / 2K · Aspect ratios 7 ratios · generate / edit
  • Docs: Qwen image API
Model launch
Claude Opus 4.6Kling V3 series
  • Model ID: kling-v3 · kling-v3-omni · kling-v3-motion-control
  • Resolution 720P / 1080P / 4K · Duration 5 / 10 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Kling video API
Model launch
Grok Imagine 1.0 · Image
  • Model ID: grok-imagine-1.0-rev · grok-imagine-1.0-edit-rev
  • Aspect ratios 1:1 / 16:9 / 9:16 / 3:2 / 2:3 · generate / edit
  • Docs: Grok Imagine image API
Grok Imagine 1.0 · Video
  • Model ID: grok-imagine-1.0-video-rev
  • Resolution 480p / 720p · Duration 6–30 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Grok Imagine video API
Model launch
Kimi K2.5
Model launch
GPT-5.2 Codex
Model launch
MiniMax M2.1
Model launch
GLM-4.7
Model launch
Gemini 3 Flash Preview
Model launch
Seedance 1.5 Pro
  • Model ID: seedance-1.5-pro
  • Resolution 480p / 720p / 1080p · Duration 5 / 10 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Seedance video API
Wan 2.6
  • Model ID: wan2.6
  • Resolution 720P / 1080P · Duration 5 / 10 / 15 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Wan video API
Model launch
GPT-5.2 series
Model launch
GPT-5.1 Codex Max
Model launch
Kling V2.6 series
  • Model ID: kling-v2-6 · kling-v2-6-motion-control
  • Resolution 720P / 1080P · Duration 5 / 10 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Kling video API
Model launch
DeepSeek V3.2Kling Video O1
  • Model ID: kling-video-o1
  • Resolution 720P / 1080P · Duration 5 / 10 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Kling video API
Model launch
Z-Image Turbo
  • Model ID: z-image-turbo
  • Resolution 1K / 2K · Aspect ratios 7 ratios
  • Docs: Z-Image API
Model launch
FLUX 2 series
  • Model ID: flux-2-pro · flux-2-flex
  • Resolution 1K / 2K · Aspect ratios 7 ratios
  • Docs: FLUX image API
Model launch
Gemini 3 Pro Image Preview
  • Model ID: gemini-3-pro-image-rev · gemini-3-pro-image
  • Resolution 1K / 2K / 4K · Aspect ratios 11 ratios · generate / edit
  • Docs: Gemini image API
Model launch
Gemini 3 Pro Preview
Model launch
GPT-5.1 series
  • Model ID: gpt-5.1 · gpt-5.1-codex · gpt-5.1-codex-mini · gpt-5.1-chat-latest
  • Docs: Text generation API
Model launch
Claude Opus 4.5
Model launch
MiniMax Hailuo 2.3
  • Model ID: minimax-hailuo-2.3 · minimax-hailuo-2.3-fast
  • Resolution 768p / 1080p · Duration 6 / 10 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Hailuo video API
Model launch
DeepSeek OCR
Model launch
Veo 3.1 series
  • Model ID: veo3.1-quality-rev · veo3.1-fast-rev · veo3.1-lite · veo3.1-quality · veo3.1-fast
  • Resolution 720p / 1080p / 4K · Duration 8 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Veo video API
GPT-5 Search API
Model launch
GPT-5 ProSora 2 Pro
  • Model ID: sora-2-pro
  • Resolution 720p / 1024p / 1080p · Duration 4–20 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Sora video API
Sora 2 Preview
  • Model ID: sora-2-preview
  • Resolution 720p · Duration 4–20 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Sora video API
Model launch
Claude Haiku 4.5
Model launch
GLM-4.6Sora 2
  • Model ID: sora-2
  • Resolution 720p · Duration 4–20 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Sora video API
Model launch
DeepSeek V3.2-Exp · Claude Sonnet 4.5
Model launch
GPT-5 CodexWan 2.5 Preview
  • Model ID: wan2.5-preview
  • Resolution 480p / 720p / 1080p · Duration 5 / 10 s · Aspect ratios 16:9 / 9:16 / 1:1
  • Docs: Wan video API
Model launch
Gemini 2.5 Flash Image Preview
  • Model ID: gemini-2.5-flash-image-preview-rev · gemini-2.5-flash-image-preview
  • Resolution 1K · Aspect ratios 11 ratios · generate / edit
  • Docs: Gemini image API
Model launch
GPT-5 series
Model launch
Gemini 2.5 Flash Lite
Model launch
Kimi K2
Model launch
Gemini 2.5 Flash / Pro
Model launch
FLUX Kontext Pro / Max
  • Model ID: flux-kontext-pro · flux-kontext-max
  • Aspect ratios 7 ratios · generate / edit
  • Docs: FLUX image API
Model launch
DeepSeek R1 (0528)
Model launch
GPT Image 1 (official)
  • Model ID: gpt-image-1
  • Aspect ratios 1:1 / 3:2 / 2:3 · generate / edit
  • Docs: GPT Image API
Model launch
o3 · o4-mini
Model launch
GPT-4.1 series
Model launch
DeepSeek V3 (0324)
Model launch
o3-mini
Model launch
HappyHorse 1.0
  • Model ID: happyhorse-1.0
  • Resolution 720P / 1080P · Duration 3–15 s · Aspect ratios 16:9 / 9:16 / 1:1 · generate / edit
  • Docs: HappyHorse video API
Model launch
o1
Model launch
GPT-4o (2024-11-20)
Model launch
GPT-4o (2024-08-06)
Model launch
GPT-4o mini
Model launch
GPT-4o
Model launch
Embeddings 3 · GPT-3.5 Turbo
  • Model ID: text-embedding-3-small · text-embedding-3-large · gpt-3.5-turbo-0125
  • Docs: Text generation API
Model launch
GPT-3.5 Turbo