Text Generation
General Chat API
Call every text model through one OpenAI-compatible endpoint — switching models means changing model
POST
An OpenAI Chat Completions-compatible endpoint and the default entry for every text model. Switching models means changing
model — the URL and key stay the same. What each model supports is in the Text Models Overview.
Authorizations
string
required
All endpoints require Bearer Token authentication.Get your API Key:Visit the API Key management page to obtain your API Key.Add it to the request header:
Body
string
required
Model ID.For the full list of models and what each supports, see the Text Models Overview; live availability comes from
GET /v1/models.Switching models means changing this field only. Which parameters work depends on the model — not every OpenAI option is available everywhere: GPT-5 and later use max_completion_tokens and reasoning_effort, Claude uses max_tokens, and gpt-5-pro, gpt-5.2-pro, gpt-5.4-pro, gpt-5.3-codex and o3-pro accept only the Responses API.array
required
List of conversation messages.
number
Controls output randomness, range 0–2.
- Lower values (e.g., 0.2) make output more deterministic.
- Higher values (e.g., 1.8) make output more random.
integer
Maximum number of tokens to generate.The maximum allowed value varies by model — refer to the specific model documentation.
boolean
Whether to use streaming output.
true: Stream the response as Server-Sent Events (SSE).false: Return the full response in one go.
number
Nucleus sampling parameter, range 0–1.Controls diversity. We recommend using either
temperature or top_p, not both.Default: 1.0number
Frequency penalty, range -2.0 to 2.0.Positive values reduce the likelihood of repeating the same words.Default: 0
number
Presence penalty, range -2.0 to 2.0.Positive values increase the likelihood of introducing new topics.Default: 0
string or array
Stop sequences.Up to 4 sequences. Generation stops when any of them is encountered.
integer
Number of completions to generate.Default: 1
integer
Output token budget (OpenAI GPT-5 and later).These models reject
max_tokens; use this field instead. Reasoning tokens come out of the same budget: finish_reason: "length" means the answer may be unfinished and must not be accepted as complete.string
Reasoning tier (OpenAI reasoning models).One of
none / low / medium / high / xhigh. GPT-5.6 and GPT-5.4 / 5.5 were accepted with none; GPT-6.1 Sol accepts low / medium / high / xhigh and rejects none; GPT-6 Astra was accepted with low only and rejects none. Do not send this field to non-OpenAI models.object
Streaming options.With
{"include_usage": true} the SSE stream emits one extra usage-only event before [DONE]. That frame’s choices can be an empty array — read usage first, then check for a text delta, or an index error will drop the cost information. See Streaming.object
Structured output.
{"type": "json_schema", "json_schema": {...}} constrains the reply to a JSON Schema; {"type": "json_object"} only guarantees valid JSON. Available only on models marked ✓ on the left of the JSON / Tools column. The whole Claude line returns 400; deepseek-v3.2-exp and deepseek-v3.1-terminus do not guarantee strict schema adherence.array
Function declarations available to the model.
[{"type": "function", "function": {"name": ..., "description": ..., "parameters": {JSON Schema}}}]. The model only generates a function name and arguments in the response’s tool_calls; the platform executes nothing — running the tool and feeding the result back is your application’s job. Available only on models marked ✓ on the right of the JSON / Tools column.string or object
How tools are selected.
"auto" (default), "none", or {"type": "function", "function": {"name": "..."}} to force a specific function. The Claude Fable family accepts only "auto"; "required" or a named function returns 400.Response
string
Unique identifier of the response.
string
Object type, always
chat.completion.integer
Creation timestamp.
string
Name of the model that actually served the request.
array
List of generated completions.
object
Token usage statistics.
string
System fingerprint (used to track backend configuration).
Available models
Every model ID, what it supports, its long-context tier and how it is billed are in the Text Models Overview. Live availability and rates come from the catalog API:Examples
Basic chat
System prompt
Multi-turn conversation
Streaming output
Usage and billing
usage.prompt_tokens is total input and already includes prompt_tokens_details.cached_tokens; usage.completion_tokens already includes completion_tokens_details.reasoning_tokens. Cached tokens are billed at the cache rate instead of the ordinary input rate, and reasoning is billed once at the output rate — neither is added twice.
usage.cost in the response is an integer quota, not dollars (500,000 quota = 1 USD). The final charge is the text ledger entry in the console billing records; pricing_pending means reconciliation is outstanding, not a zero charge. Full rules, ledger states and the cost formula are in the Text Models Overview · Billing rules.
Not supported
The following return400 and are not billed:
- Vendor built-in tools — any entry in
toolswhose type is notfunctionorcustom. Client-executed function tools work normally. web_search_options.service_tierset to anything other thanstandard/default.
tools, tool_choice or response_format to a model that does not support tool calling or structured output also returns 400. What each model supports is in the Text Models Overview. Models billed as input + output ignore explicit cache_control and bill all input at the input rate; how other models handle it is in Text Models Overview · Caching.