Gemini Image Series
Nano Banana 2.1
Google Nano Banana 2.1 official line: text-to-image and multi-reference editing, up to 4 images per request, billed from actual token usage
POST
Nano Banana 2.1
Nano Banana 2.1 (
A successful submission returns a
See Submit a task for
The example above is the result of
gemini-nano-banana-2.1) is Google’s image model for text-to-image and editing with up to 14 reference images. It outputs at 1k, 2k and 4k, returns up to 4 images per request, and is billed from actual token usage. It is a different model from Nano Banana 2 (gemini-3.1-flash-image, see Nano Banana Official): there is no 0.5k tier and no google_search or google_image_search, so drop those when migrating from Nano Banana 2. When the price of an image has to be known before submission, use Nano Banana 2.1 Economy.
The official and economy lines never switch automatically: the model ID selects the line.
Quick start
task_id. Retrieve the result with GET /v1/tasks/{task_id}, wait inline with Prefer: wait, or configure a webhook.
Request parameters
string
required
Must be
gemini-nano-banana-2.1.string
default:"generate"
generate— generate from textedit— rewrite from reference images;image_urlsis required
edit; the only difference is how many entries image_urls carries.string
required
Image description or editing instruction, in English or Chinese.
string[]
Reference image URLs. Required for
edit, up to 14. Reference images are billed as image input — see “Pricing”.string
default:"auto"
Frame ratio: 14 fixed ratios plus
auto — auto, 1:1, 2:3, 3:2, 3:4, 4:3, 4:5, 5:4, 9:16, 16:9, 21:9, 1:4, 4:1, 1:8, 8:1.1:4, 4:1, 1:8 and 8:1 suit long horizontal and vertical banners and are available at 1k, 2k and 4k.auto lets the model decide: it picks from the prompt for text-to-image, and follows the input image for edits. When omitted, auto is used.string
default:"1k"
Output resolution:
1k, 2k or 4k. 0.5k is not supported and returns 400.integer
default:"1"
Number of images to generate,
1–4. Each image is billed for its own output image tokens, so a larger n costs more.callback_url, callback_events, Prefer: wait, Idempotency-Key and the maximum-cost header.
Limits
Pricing
Billed from actual token usage across four dimensions:
More reference images, a higher resolution and a larger
n all raise the cost of a call.
Rates are in
price_config on GET /v1/models and in the console’s Model Market.What a single call actually cost is the cost field on the task response: an integer quota at 500,000 quota = 1 USD, not dollars.Holds and settlement
Submitting a task places a hold for the estimated usage; the task then settles from actual usage and the difference is released. The settled amount never exceeds the hold. The estimate is built from:- Image output — expected output tokens per image, looked up by
resolution, multiplied byn - Text output — reserved at the model’s maximum output tokens, once per request
- Text input — reserved by the prompt’s byte length, with a floor of 64
- Reference images — 768px tiles at 258 tokens per tile
price_config.image_usage_reservation on GET /v1/models.
You are only billed for a successfully generated image. Failed and cancelled tasks, and tasks that return no usable image, are refunded in full.
Response
Completed task
n: 2. result.images has one entry per image, where url is an array and expires_at is when the link stops working. See Query Task Status for the full field list.