Skip to main content
POST
The OpenAI Responses protocol endpoint. Every text model except deepseek-v3.1-terminus can be called here. gpt-5-pro, gpt-5.2-pro, gpt-5.4-pro, gpt-5.3-codex and o3-pro accept only this endpoint, not Chat Completions.

Authorizations

string
required
All endpoints require Bearer Token authentication.Get your API Key:Visit the API Key management page to obtain your API Key.Add it to the request header:

Body

string
required
Model ID.Every text model except deepseek-v3.1-terminus can be called on this endpoint. Five models accept only this endpoint:
  • gpt-5-pro
  • gpt-5.2-pro
  • gpt-5.4-pro
  • gpt-5.3-codex
  • o3-pro
These five return 400 on /v1/chat/completions; deepseek-v3.1-terminus returns 400 here. How each model is billed is in the Text Models Overview.
string or array
required
Input content — a string or an array of messages.A string is a one-shot text input; the array form carries multiple turns:
integer
Output budget for this request.Reasoning tokens come out of the same budget. status: "incomplete" means the answer may be unfinished and must not be accepted as complete.
array
Tool list.
These five models do not support tool calling or structured output, so tools, tool_choice and response_format return 400 and are not billed. For function calling, use another model that supports tools.
number
Controls output randomness, range 0–2.Default: 1.0
integer
Maximum number of tokens to generate.
boolean
Whether to use streaming output.Default: false

Response

string
Unique identifier of the response.
string
Object type, always response.
integer
Creation timestamp.
string
Name of the model that actually served the request (e.g., gpt-5-2025-08-07).
string
Response status.Possible values:
  • completed — Done
  • in_progress — Processing
  • failed — Failed
  • cancelled — Cancelled
array
Output content array.
object
Token usage statistics.
object
Reasoning configuration (thinking models only).
number
The sampling temperature actually used.
number
The nucleus sampling parameter actually used.
string
Tool selection strategy.
array
List of tools used.
boolean
Whether parallel tool calls are allowed.
boolean
Whether the conversation history is stored.
string
Service tier.
string
Truncation strategy.
object
Text format configuration.
boolean
Whether this is a background task.
object
Error info, if any.
object
Metadata.

Examples

Single input

Multi-turn input

Capping the output budget

Reasoning tokens come out of max_output_tokens as well. A response with status: "incomplete" means the budget ran out and the output may be cut short.

Streaming

Usage and billing

usage.input_tokens is total input (including input_tokens_details.cached_tokens), and usage.output_tokens already includes output_tokens_details.reasoning_tokens — reasoning is billed once, at the output rate. All five models bill input and output only, with no separate cache rate; any cache statistics in the response bill at the ordinary input rate. Rates and the shared rules are in the Text Models Overview · Billing rules. Streaming bills from the terminal usage: take usage from the response.completed, response.incomplete, response.failed or response.cancelled event. A failed or truncated response can still incur token charges; HTTP 200, partial text or [DONE] do not prove you received complete usage.
If no terminal usage arrives, the held credit stays pending reconciliation. Check that call’s usage record in the console before resubmitting; never retry automatically.

Not supported

The following return 400 and are not billed:
  • Hosted tools — file_search, remote_mcp and other vendor built-ins; for which models support web_search and how it is billed, see Text Models Overview · Billing rules.
  • Tool calling and structured output on the five models above — see the tools note above.
  • service_tier set to anything other than standard / default.
  • Image and video input.