Skip to main content
POST
MAI-Image-2.6
MAI-Image-2.6 is Microsoft’s image model for text-to-image, single-image editing and multi-image composition from up to 5 reference images, with PNG output. This page covers the standard mai-image-2.6 and the faster, lower-cost mai-image-2.6-flash: the two share the same actions, parameters and billing model and differ only in rates — switching models means changing model and nothing else. For masked inpainting use gpt-image-2.

Quick start

A successful submission returns a task_id. Retrieve the result with GET /v1/tasks/{task_id}, wait inline with Prefer: wait, or configure a webhook.

Request parameters

string
required
One of mai-image-2.6 or mai-image-2.6-flash.
string
default:"generate"
  • generate — generate from text; takes no reference images
  • edit — edit from reference images; image_urls is required. One image is a single-image edit, 2–5 images is multi-image composition
For edits the model sets the output size, about 1 megapixel and close to the reference ratio: aspect_ratio and resolution have no effect, and width and height cannot be sent.
string
required
Image description or editing instruction, in English or Chinese.
string[]
Reference image URLs. Required for edit, up to 5; cannot be sent with generate. Reference images are billed as image input tokens — see “Pricing”.
string
default:"1:1"
Aspect ratio: auto, 1:1, 4:3, 3:4, 3:2, 2:3, 16:9, 9:16, 2:1, 1:2, 21:9, 9:21, 4:1, 1:4. auto lets the model pick the ratio from the prompt.Applies to generate only; when width and height are sent, the pixel size wins.
string
default:"1k"
Output tier: 1k or 2k. 1k is about 1 megapixel and 2k about 2.3 megapixels; the short side is at least 768, so at extreme ratios 1k goes above 1 megapixel. Output sizes per aspect ratio are listed under “Pricing”.Applies to generate only; when width and height are sent, it plays no part in the size.
integer
default:"1"
Number of images to generate. Value: 1. Each request returns one image.
boolean
default:"false"
Search for real-time information before generating; useful for real people, places and events. Available for both generate and edit.
integer
Exact output width in pixels, sent together with height; at least 768, width × height at most 2,359,296 (the pixel count of 1536 × 1536), rounded down to a multiple of 32. Overrides aspect_ratio and resolution. action: "generate" only; sending it with edit returns 400.The cap is on total pixels, not on each side: 2048 × 1152 and 3072 × 768 are both valid. Values that are not multiples of 32 are rounded down, so 1000 × 1000 comes out as 992 × 992.
integer
Exact output height in pixels, sent together with width; same limits as width. action: "generate" only.
See Submit Task for callback_url, callback_events, Prefer: wait, Idempotency-Key, and the maximum-cost header. Neither model accepts seed, quality, background, output_format, negative_prompt or mask_url: output is always PNG and masked editing is not supported. These and any other parameters not listed above return 400 and are not billed.

Limits

Pricing

Billed from actual token usage across three dimensions: text input, image input and image output. Both models use the same dimensions at different rates; every rate on mai-image-2.6-flash is lower than on mai-image-2.6.
  • Image output — output image tokens = output width × height ÷ 1024. For example, 1024×1024 is 1,024 tokens and 1536×1536 is 2,304 tokens
  • Image input — each reference image is about width × height ÷ 1024 tokens
  • Text input — the prompt’s tokens
Output tokens depend only on the pixel count of the output, so aspect ratio, tier and width / height all affect the cost; an edit outputs about 1 megapixel, roughly 1,000 tokens.

Output sizes

Text-to-image produces these sizes (width × height, pixels) for each aspect_ratio and resolution. With width and height, the output uses the values sent, rounded down to multiples of 32. Because the short side is at least 768, the more extreme the ratio, the more pixels the 1k tier produces: 4:1 outputs 3072×768 at both 1k and 2k, the same output tokens as 1:1 at 2k. With auto, the size follows the ratio the model picks.

Holds and settlement

Submitting a task places a hold for the estimated usage; the task then settles from actual usage and the difference is released. The estimate is built from:
  • Image output — 2,304 tokens, the largest output size, regardless of aspect ratio and tier
  • Text input — a floor of 64 tokens
  • Reference images — width × height ÷ 1024 tokens each; 4,096 tokens when the size cannot be read
Because image output is reserved at the maximum, the hold is above the final settled amount whenever the output is smaller (for example 1:1 at 1k, or an edit), and the difference is released when the task finishes.
Per-model rates are in price_config on GET /v1/models, and in the console’s Model Market.What a single call actually cost is the cost field on the task response — an integer quota at 500,000 quota = 1 USD.
You are only billed for a successfully generated image. Failed and cancelled tasks, and tasks that return no usable image, are refunded in full. The final charge is the cost field on the task response — an integer quota, not dollars.

Response

Completed task
Generated images (PNG) come back in result.images, where url is an array and expires_at is when the link stops working. See Query Task Status for the full field list.

Available models