Image Generation
MAI-Image-2.6
Microsoft MAI-Image-2.6 and its Flash version: text-to-image, single-image editing and composition from up to 5 references, exact pixel sizes, billed from actual token usage
POST
MAI-Image-2.6
MAI-Image-2.6 is Microsoft’s image model for text-to-image, single-image editing and multi-image composition from up to 5 reference images, with PNG output. This page covers the standard
A successful submission returns a
See Submit Task for
You are only billed for a successfully generated image. Failed and cancelled tasks, and tasks that return no usable image, are refunded in full. The final charge is the
Generated images (PNG) come back in
mai-image-2.6 and the faster, lower-cost mai-image-2.6-flash: the two share the same actions, parameters and billing model and differ only in rates — switching models means changing model and nothing else. For masked inpainting use gpt-image-2.
Quick start
task_id. Retrieve the result with GET /v1/tasks/{task_id}, wait inline with Prefer: wait, or configure a webhook.
Request parameters
string
required
One of
mai-image-2.6 or mai-image-2.6-flash.string
default:"generate"
generate— generate from text; takes no reference imagesedit— edit from reference images;image_urlsis required. One image is a single-image edit, 2–5 images is multi-image composition
aspect_ratio and resolution have no effect, and width and height cannot be sent.string
required
Image description or editing instruction, in English or Chinese.
string[]
Reference image URLs. Required for
edit, up to 5; cannot be sent with generate. Reference images are billed as image input tokens — see “Pricing”.string
default:"1:1"
Aspect ratio:
auto, 1:1, 4:3, 3:4, 3:2, 2:3, 16:9, 9:16, 2:1, 1:2, 21:9, 9:21, 4:1, 1:4. auto lets the model pick the ratio from the prompt.Applies to generate only; when width and height are sent, the pixel size wins.string
default:"1k"
Output tier:
1k or 2k. 1k is about 1 megapixel and 2k about 2.3 megapixels; the short side is at least 768, so at extreme ratios 1k goes above 1 megapixel. Output sizes per aspect ratio are listed under “Pricing”.Applies to generate only; when width and height are sent, it plays no part in the size.integer
default:"1"
Number of images to generate. Value:
1. Each request returns one image.boolean
default:"false"
Search for real-time information before generating; useful for real people, places and events. Available for both
generate and edit.integer
Exact output width in pixels, sent together with
height; at least 768, width × height at most 2,359,296 (the pixel count of 1536 × 1536), rounded down to a multiple of 32. Overrides aspect_ratio and resolution. action: "generate" only; sending it with edit returns 400.The cap is on total pixels, not on each side: 2048 × 1152 and 3072 × 768 are both valid. Values that are not multiples of 32 are rounded down, so 1000 × 1000 comes out as 992 × 992.integer
Exact output height in pixels, sent together with
width; same limits as width. action: "generate" only.callback_url, callback_events, Prefer: wait, Idempotency-Key, and the maximum-cost header.
Neither model accepts seed, quality, background, output_format, negative_prompt or mask_url: output is always PNG and masked editing is not supported. These and any other parameters not listed above return 400 and are not billed.
Limits
Pricing
Billed from actual token usage across three dimensions: text input, image input and image output. Both models use the same dimensions at different rates; every rate onmai-image-2.6-flash is lower than on mai-image-2.6.
- Image output — output image tokens = output width × height ÷ 1024. For example, 1024×1024 is 1,024 tokens and 1536×1536 is 2,304 tokens
- Image input — each reference image is about width × height ÷ 1024 tokens
- Text input — the prompt’s tokens
width / height all affect the cost; an edit outputs about 1 megapixel, roughly 1,000 tokens.
Output sizes
Text-to-image produces these sizes (width × height, pixels) for eachaspect_ratio and resolution. With width and height, the output uses the values sent, rounded down to multiples of 32.
Because the short side is at least 768, the more extreme the ratio, the more pixels the
1k tier produces: 4:1 outputs 3072×768 at both 1k and 2k, the same output tokens as 1:1 at 2k. With auto, the size follows the ratio the model picks.
Holds and settlement
Submitting a task places a hold for the estimated usage; the task then settles from actual usage and the difference is released. The estimate is built from:- Image output — 2,304 tokens, the largest output size, regardless of aspect ratio and tier
- Text input — a floor of 64 tokens
- Reference images — width × height ÷ 1024 tokens each; 4,096 tokens when the size cannot be read
1:1 at 1k, or an edit), and the difference is released when the task finishes.
Per-model rates are in
price_config on GET /v1/models, and in the console’s Model Market.What a single call actually cost is the cost field on the task response — an integer quota at 500,000 quota = 1 USD.cost field on the task response — an integer quota, not dollars.
Response
Completed task
result.images, where url is an array and expires_at is when the link stops working. See Query Task Status for the full field list.