Skip to main content
POST
A synchronous endpoint: the response is the audio file itself, with no task_id and no polling. It serves gpt-4o-mini-tts only; for songs and sung vocals, use Suno Music Generation.

Request parameters

string
required
Model ID: gpt-4o-mini-tts.
string
required
The text to convert, up to 4,096 characters.
string
required
Voice, six options:
  • alloy: neutral, balanced
  • echo: male, calm
  • fable: British, narrative
  • onyx: male, deep
  • nova: female, energetic
  • shimmer: female, gentle
string
default:"wav"
Audio format: wav (uncompressed), opus (streaming), aac, flac (lossless), pcm (raw audio data).
number
default:"1.0"
Playback speed, value 1.0.
Anything beyond the fields above returns 400 and is not billed.

Limits

Pricing

Billed by the number of input text characters, with no output charge. Audio length, voice and response_format make no difference — the character count is the only billing dimension, so the cost of a call is known in advance. Rates are in price_config on GET /v1/models and in the console’s Model Market. The billing unit is quota, at 500,000 quota = 1 USD, not dollars. A failed request is not billed.

Response

The response body is the audio file as a binary stream, in the format given by response_format (wav by default) — save it directly to a file. On error the response is a JSON error object and carries no audio.

Available models

Not supported

  • instructions, for setting tone, pace and emotion in plain language
  • Streaming audio responses
  • The mp3 output format
These requests return 400 and are not billed.