Audio
Text-to-Speech
Convert text into natural, fluent speech
POST
A synchronous endpoint: the response is the audio file itself, with no
Anything beyond the fields above returns
task_id and no polling. It serves gpt-4o-mini-tts only; for songs and sung vocals, use Suno Music Generation.
Request parameters
string
required
Model ID:
gpt-4o-mini-tts.string
required
The text to convert, up to 4,096 characters.
string
required
Voice, six options:
alloy: neutral, balancedecho: male, calmfable: British, narrativeonyx: male, deepnova: female, energeticshimmer: female, gentle
string
default:"wav"
Audio format:
wav (uncompressed), opus (streaming), aac, flac (lossless), pcm (raw audio data).number
default:"1.0"
Playback speed, value
1.0.400 and is not billed.
Limits
Pricing
Billed by the number of input text characters, with no output charge. Audio length,voice and response_format make no difference — the character count is the only billing dimension, so the cost of a call is known in advance.
Rates are in price_config on GET /v1/models and in the console’s Model Market. The billing unit is quota, at 500,000 quota = 1 USD, not dollars.
A failed request is not billed.
Response
The response body is the audio file as a binary stream, in the format given byresponse_format (wav by default) — save it directly to a file. On error the response is a JSON error object and carries no audio.
Available models
Not supported
instructions, for setting tone, pace and emotion in plain language- Streaming audio responses
- The
mp3output format
400 and are not billed.