Request Format
All requests use JSON and require these headers:Successful Responses
Response shapes vary by endpoint, each following its own API spec:- Chat Completions — OpenAI standard response format
- Claude Messages — Anthropic standard response format
- Gemini — Google standard response format
- Asynchronous Tasks — QWave API unified task response format
Billing Field (usage.cost)
The OpenAI-format responses from/v1/chat/completions and /v1/embeddings return the call’s billed quota in usage.cost. Its unit is quota (an integer), not US dollars (USD). For example, cost: 150 means 150 quota.
- Non-streaming Chat Completions and Embeddings →
usage.costin the response body - Streaming Chat Completions (SSE) → the last data frame carrying
usage(see Streaming)
/v1/messages), OpenAI Responses (/v1/responses), and native Gemini endpoints use their own usage structures. They do not currently define a shared gateway usage.cost field. Do not interpret an upstream field with the same name as gateway quota; check these calls’ charges in the console usage records.
Async tasks have separate cost and settlement-status fields; see Query Task. Do not add task cost fields directly to text usage.cost as though their units were the same.
Error Responses
Error structures differ slightly across endpoints: Synchronous endpoints (chat / messages / images, etc.) — OpenAI-compatible format with atype field:
/v1/tasks/...) — Simplified format with only message + code: