Skip to main content

Request Format

All requests use JSON and require these headers:

Successful Responses

Response shapes vary by endpoint, each following its own API spec:
  • Chat Completions — OpenAI standard response format
  • Claude Messages — Anthropic standard response format
  • Gemini — Google standard response format
  • Asynchronous Tasks — QWave API unified task response format

Billing Field (usage.cost)

The OpenAI-format responses from /v1/chat/completions and /v1/embeddings return the call’s billed quota in usage.cost. Its unit is quota (an integer), not US dollars (USD). For example, cost: 150 means 150 quota.
  • Non-streaming Chat Completions and Embeddings → usage.cost in the response body
  • Streaming Chat Completions (SSE) → the last data frame carrying usage (see Streaming)
Claude Messages (/v1/messages), OpenAI Responses (/v1/responses), and native Gemini endpoints use their own usage structures. They do not currently define a shared gateway usage.cost field. Do not interpret an upstream field with the same name as gateway quota; check these calls’ charges in the console usage records. Async tasks have separate cost and settlement-status fields; see Query Task. Do not add task cost fields directly to text usage.cost as though their units were the same.

Error Responses

Error structures differ slightly across endpoints: Synchronous endpoints (chat / messages / images, etc.) — OpenAI-compatible format with a type field:
Asynchronous task endpoints (/v1/tasks/...) — Simplified format with only message + code:

Status Codes