Image Generation
FLUX 3 Image
Black Forest Labs FLUX 3 Image: text-to-image, single-image editing and multi-reference with up to 10 images, bbox layout control in the prompt, billed per image by output tier
POST
FLUX 3 Image
FLUX 3 Image is Black Forest Labs’ third-generation image model. One model covers text-to-image, single-image editing and multi-image reference with up to 10 images, at five output tiers from
A successful submission returns a
See Submit Task for
You are only billed for a successfully generated image. Failed and cancelled tasks, and tasks that return no usable image, are refunded in full. The final charge is the
Generated images come back in
To move an element, point
768sq to 4k. Tagging elements in the prompt and appending bounding boxes lets you set the layout or change only part of the image — see “Examples”. For FLUX.2 and FLUX.1 Kontext, see FLUX Series.
Quick start
task_id. Retrieve the result with GET /v1/tasks/{task_id}, wait inline with Prefer: wait, or configure a webhook.
Request parameters
string
required
Must be
flux-3-image.string
default:"generate"
generate— generate from text; takes no reference imagesedit— edit from reference images;image_urlsis required. One image is a single-image edit, 2–10 images is multi-image reference
string
required
Image description or editing instruction. Negative prompts are not supported; describe what you want to see.You can refer to elements with
<tags> and append a bbox array to the same string to set the layout or the region to edit — see “Examples”.string[]
Reference image URLs. Required for
edit, up to 10; cannot be sent with generate. References are numbered in order and can be referred to in the prompt as ref_image_0, ref_image_1, … or Image 1, Image 2, …. Reference images are not billed.string
default:"auto"
Aspect ratio:
auto, 21:9, 2:1, 16:9, 3:2, 7:5, 4:3, 5:4, 1:1, 4:5, 3:4, 5:7, 2:3, 9:16, 1:2, 9:21.With auto, edit follows the first reference image, and generate decides from the prompt, falling back to 1:1.string
default:"1k"
Output tier, which is also the billing tier:
768sq (about 768×768 square), 1k (about 1 MP), 1.5k (about 2 MP), 2k (about 4 MP), 4k (about 16 MP). The actual pixel size is that of the returned image. 4k can take several minutes.integer
default:"1"
Number of images to generate. Value:
1. Each request returns one image.integer
default:"2"
Content safety tolerance,
0–4; 0 is the strictest and higher values are more permissive.boolean
default:"true"
Allow web or image search before generating. Set
false to turn it off.callback_url, callback_events, Prefer: wait, Idempotency-Key, and the maximum-cost header.
This model does not accept seed, steps, guidance, output_format, prompt_upsampling, negative_prompt or mask_url, nor pixel sizes such as width and height: pick the resolution with resolution and the frame with aspect_ratio. These and any other parameters not listed above return 400 and are not billed.
Limits
Pricing
Billed at a fixed price per generated image, withresolution as the only billing dimension: 768sq, 1k, 1.5k, 2k and 4k are one tier each, and higher tiers cost more. generate and edit cost the same, reference images are not billed, and aspect ratio, prompt length, grounding and safety_tolerance do not affect the price, so the cost of a request is known before you submit it. Each request returns one image and is billed as one image.
Without resolution, the default 1k tier is billed. The default entry in price_config.image_prices is the fallback price and equals the 1k tier.
Rates are in
price_config on GET /v1/models, and in the console’s Model Market.What a single call actually cost is the cost field on the task response — an integer quota at 500,000 quota = 1 USD.cost field on the task response — an integer quota, not dollars.
Response
Completed task
result.images, where url is an array and expires_at is when the link stops working. See Query Task Status for the full field list.
Examples
Layout and local edits are written intoprompt; they are not separate request parameters. Write the instruction in natural language, refer to elements with <tags> (such as <car_1>), then append a JSON array to the same string, one object per box.
Every box is written as
[y1, x1, y2, x2] (top, left, bottom, right) in coordinates normalized to 0–1000: the top-left corner is [0, 0] and the bottom-right is [1000, 1000]. They are not pixels.
Local edit
Turn the car inside the box red and keep the background unchanged. The request body below is submitted the same way as in “Quick start”:from at the reference image, put the original position in src_bbox and the new position in tgt_bbox.
Text-to-image layout
Without reference images, each box uses three fields:id, bbox and desc. The coordinate grid stretches with the frame, so send aspect_ratio explicitly:
- The bbox array is part of the
promptstring. In hand-written JSON, escape the inner double quotes as\"; SDKs and JSON serializers do this for you. - Tags in the prompt map one-to-one to
idvalues in the array; identifiers such asref_image_0point to the input reference images. - List the regions to keep as well, and state what to keep in
desc. - Local edits are done with bboxes; this model does not accept
mask_url.