Skip to main content
POST
The native Anthropic Messages endpoint, kept for apps already built on the Anthropic SDK. Every text model except the GPT-5.6 / GPT-6 / GPT-6.1 series and the Responses-only models can be called here. Apart from web search (see Text Models Overview · Billing rules), it exposes nothing beyond what Chat offers, so new integrations should use the General Chat API.

Authorizations

string
API key used for authentication. Send either this header or Authorization.Visit the API Key management page to obtain your API Key.Add it to the request header:
string
API key used for authentication. Send either this header or x-api-key.
string
required
API version.Specifies which Claude API version to use.Example: 2023-06-01
Set the Anthropic SDK base URL to https://api.qingbo.ai; the SDK appends /v1/messages itself, so /v1 does not need to be added again.

Body

string
required
Model ID.Every text model can be called on this endpoint except the GPT-5.6, GPT-6 and GPT-6.1 series (gpt-5.6-luna, gpt-5.6-sol, gpt-5.6-terra, gpt-6-astra, gpt-6-luna, gpt-6-sol, gpt-6.1-sol) and the models that accept only the Responses API: gpt-5-pro, gpt-5.2-pro, gpt-5.4-pro, gpt-5.3-codex, o3-pro. These models return 400 here.Cache tiers and supported capabilities for each model are in the Text Models Overview.Image input is outside the accepted scope.
array
required
Message list with alternating user and assistant roles.
integer
required
Maximum tokens to generate.Maximum number of tokens before generation stops. The model may stop earlier.Maximum value varies by model — refer to the model documentation.Minimum: 1
string | array
System prompt.The system prompt defines Claude’s role, personality, goals, and instructions.String format:
Structured format:
number
Temperature, range 0–1.Controls output randomness:
  • Low values (e.g., 0.2): more deterministic, more conservative
  • High values (e.g., 0.8): more random, more creative
Default: 1.0
number
Nucleus sampling parameter, range 0–1.Uses nucleus sampling. We recommend using either temperature or top_p, not both.Default: 1.0
integer
Top-K sampling.Sample only from the top K highest-probability options to remove “long-tail” low-probability responses.Recommended only for advanced use cases.
boolean
Whether to enable streaming output.
  • true: Stream the response progressively via Server-Sent Events (SSE).
  • false: Return the full response in one go.
Default: false
array
Stop sequences.Custom text sequences that stop generation when encountered. Up to 4 sequences, each up to 32 tokens long.
object
Metadata.An object used to track or identify the request.

Response

string
Unique identifier of the message.
string
Object type, always message.
string
Role, always assistant.
array
Array of message content.
string
Name of the model that actually served the request.
string
Reason generation stopped.Possible values:
  • end_turn — Natural completion
  • max_tokens — Reached max token limit
  • stop_sequence — Stop sequence encountered
  • tool_use — Tool use
string
The triggering stop sequence (if any).
object
Token usage statistics.

Examples

Single turn

Multi-turn conversation

Using a system prompt

Prefilling the response

Streaming

Set stream: true on the same URL. Read message_start, then the content block events, then message_delta (which carries the actual usage and stop_reason), and finally message_stop. The initial usage in message_start is not the final bill and must not be billed against.

Usage and billing

Native input_tokens differs from Chat’s prompt_tokens: the native field excludes cache reads and cache writes, so total input is the sum of all three. output_tokens already includes reasoning tokens. Native responses carry no usage.cost field; check the console usage record for the actual charge. Billing rules are in the Text Models Overview · Billing rules.

Not supported

  • Structured output — Claude models do not accept response_format with json_schema or json_object.
  • Tool limits — tools / tool_choice on a model without tool calling (see the JSON / Tools column in the Text Models Overview); tool_choice of any or tool on the Fable models, claude-opus-5-5 and claude-sonnet-5-5 (they accept only auto).
  • Image input.
These return 400 and are not billed. Explicit cache_control does work here: put the reusable long prefix in the system array — request shape and hit conditions are in the Text Models Overview · Caching.