Text Generation
Responses API
The OpenAI Responses protocol endpoint, available for every text model except deepseek-v3.1-terminus
POST
The OpenAI Responses protocol endpoint. Every text model except
Reasoning tokens come out of
deepseek-v3.1-terminus can be called here. gpt-5-pro, gpt-5.2-pro, gpt-5.4-pro, gpt-5.3-codex and o3-pro accept only this endpoint, not Chat Completions.
Authorizations
string
required
All endpoints require Bearer Token authentication.Get your API Key:Visit the API Key management page to obtain your API Key.Add it to the request header:
Body
string
required
Model ID.Every text model except
deepseek-v3.1-terminus can be called on this endpoint. Five models accept only this endpoint:gpt-5-progpt-5.2-progpt-5.4-progpt-5.3-codexo3-pro
These five return
400 on /v1/chat/completions; deepseek-v3.1-terminus returns 400 here. How each model is billed is in the Text Models Overview.string or array
required
Input content — a string or an array of messages.A string is a one-shot text input; the array form carries multiple turns:
integer
Output budget for this request.Reasoning tokens come out of the same budget.
status: "incomplete" means the answer may be unfinished and must not be accepted as complete.array
Tool list.
number
Controls output randomness, range 0–2.Default: 1.0
integer
Maximum number of tokens to generate.
boolean
Whether to use streaming output.Default: false
Response
string
Unique identifier of the response.
string
Object type, always
response.integer
Creation timestamp.
string
Name of the model that actually served the request (e.g.,
gpt-5-2025-08-07).string
Response status.Possible values:
completed— Donein_progress— Processingfailed— Failedcancelled— Cancelled
array
Output content array.
object
Token usage statistics.
object
Reasoning configuration (thinking models only).
number
The sampling temperature actually used.
number
The nucleus sampling parameter actually used.
string
Tool selection strategy.
array
List of tools used.
boolean
Whether parallel tool calls are allowed.
boolean
Whether the conversation history is stored.
string
Service tier.
string
Truncation strategy.
object
Text format configuration.
boolean
Whether this is a background task.
object
Error info, if any.
object
Metadata.
Examples
Single input
Multi-turn input
Capping the output budget
max_output_tokens as well. A response with status: "incomplete" means the budget ran out and the output may be cut short.
Streaming
Usage and billing
usage.input_tokens is total input (including input_tokens_details.cached_tokens), and usage.output_tokens already includes output_tokens_details.reasoning_tokens — reasoning is billed once, at the output rate. All five models bill input and output only, with no separate cache rate; any cache statistics in the response bill at the ordinary input rate. Rates and the shared rules are in the Text Models Overview · Billing rules.
Streaming bills from the terminal usage: take usage from the response.completed, response.incomplete, response.failed or response.cancelled event. A failed or truncated response can still incur token charges; HTTP 200, partial text or [DONE] do not prove you received complete usage.
Not supported
The following return400 and are not billed:
- Hosted tools —
file_search,remote_mcpand other vendor built-ins; for which models supportweb_searchand how it is billed, see Text Models Overview · Billing rules. - Tool calling and structured output on the five models above — see the
toolsnote above. service_tierset to anything other thanstandard/default.- Image and video input.