Skip to main content
POST
Upstream shutdown notice: OpenAI has announced that whisper-1, gpt-4o-transcribe and gpt-4o-mini-transcribe shut down on 2027-02-26; they are delisted here the same day. Replacement models will be announced in Service Notices ahead of time. Keep model configurable in new integrations.
A synchronous endpoint: upload an audio file and the request returns the transcript, with no task_id and no polling. All three models share this endpoint — switching between them means changing model only. For the opposite direction, reading text aloud, use Text-to-Speech.

Request parameters

The request body uses multipart/form-data.
file
required
The audio file to transcribe. Supported formats: mp3, mp4, mpeg, mpga, m4a, wav, webm.
string
required
Model ID: gpt-4o-transcribe, gpt-4o-mini-transcribe or whisper-1.
Anything beyond the fields above returns 400 and is not billed.

Pricing

Billed by audio length in minutes. The three models have different rates; transcript length, spoken language and file format make no difference to the price. Rates are in price_config on GET /v1/models and in the console’s Model Market. The billing unit is quota, at 500,000 quota = 1 USD, not dollars. A failed request is not billed.

Response

string
The transcribed text.

Available models