Text Generation
Claude Messages API
The native Messages endpoint, kept for apps already built on the Anthropic SDK
POST
The native Anthropic Messages endpoint, kept for apps already built on the Anthropic SDK. Every text model except the GPT-5.6 / GPT-6 / GPT-6.1 series and the Responses-only models can be called here. Apart from web search (see Text Models Overview · Billing rules), it exposes nothing beyond what Chat offers, so new integrations should use the General Chat API.
Set the Anthropic SDK base URL to
Authorizations
string
API key used for authentication. Send either this header or
Authorization.Visit the API Key management page to obtain your API Key.Add it to the request header:string
API key used for authentication. Send either this header or
x-api-key.string
required
API version.Specifies which Claude API version to use.Example:
2023-06-01https://api.qingbo.ai; the SDK appends /v1/messages itself, so /v1 does not need to be added again.
Body
string
required
Model ID.Every text model can be called on this endpoint except the GPT-5.6, GPT-6 and GPT-6.1 series (
gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, gpt-6-luna, gpt-6-sol, gpt-6.1-sol) and the models that accept only the Responses API: gpt-5-pro, gpt-5.2-pro, gpt-5.4-pro, gpt-5.3-codex, o3-pro. These models return 400 here.Cache tiers and supported capabilities for each model are in the Text Models Overview.Image input is outside the accepted scope.array
required
Message list with alternating
user and assistant roles.integer
required
Maximum tokens to generate.Maximum number of tokens before generation stops. The model may stop earlier.Maximum value varies by model — refer to the model documentation.Minimum: 1
string | array
System prompt.The system prompt defines Claude’s role, personality, goals, and instructions.String format:Structured format:
number
Temperature, range 0–1.Controls output randomness:
- Low values (e.g., 0.2): more deterministic, more conservative
- High values (e.g., 0.8): more random, more creative
number
Nucleus sampling parameter, range 0–1.Uses nucleus sampling. We recommend using either
temperature or top_p, not both.Default: 1.0integer
Top-K sampling.Sample only from the top K highest-probability options to remove “long-tail” low-probability responses.Recommended only for advanced use cases.
boolean
Whether to enable streaming output.
true: Stream the response progressively via Server-Sent Events (SSE).false: Return the full response in one go.
array
Stop sequences.Custom text sequences that stop generation when encountered. Up to 4 sequences, each up to 32 tokens long.
object
Metadata.An object used to track or identify the request.
Response
string
Unique identifier of the message.
string
Object type, always
message.string
Role, always
assistant.array
Array of message content.
string
Name of the model that actually served the request.
string
Reason generation stopped.Possible values:
end_turn— Natural completionmax_tokens— Reached max token limitstop_sequence— Stop sequence encounteredtool_use— Tool use
string
The triggering stop sequence (if any).
object
Token usage statistics.
Examples
Single turn
Multi-turn conversation
Using a system prompt
Prefilling the response
Streaming
Setstream: true on the same URL. Read message_start, then the content block events, then message_delta (which carries the actual usage and stop_reason), and finally message_stop. The initial usage in message_start is not the final bill and must not be billed against.
Usage and billing
Nativeinput_tokens differs from Chat’s prompt_tokens: the native field excludes cache reads and cache writes, so total input is the sum of all three. output_tokens already includes reasoning tokens. Native responses carry no usage.cost field; check the console usage record for the actual charge. Billing rules are in the Text Models Overview · Billing rules.
Not supported
- Structured output — Claude models do not accept
response_formatwithjson_schemaorjson_object. - Tool limits —
tools/tool_choiceon a model without tool calling (see the JSON / Tools column in the Text Models Overview);tool_choiceofanyortoolon the Fable models,claude-opus-5-5andclaude-sonnet-5-5(they accept onlyauto). - Image input.
400 and are not billed. Explicit cache_control does work here: put the reusable long prefix in the system array — request shape and hit conditions are in the Text Models Overview · Caching.