Error format
When a /v1 endpoint fails it answers with a 4xx / 5xx status and JSON like the example below: message is for people, type is the broad class, code is a stable machine-readable identifier and param names the field at fault (null when there is none). Branch on the HTTP status and code, never on the message text: it may be English, Chinese or both, and its wording changes.
HTTP 402 · /v1
{
"error": {
"message": "Insufficient balance — please top up",
"type": "insufficient_quota",
"code": "insufficient_quota",
"param": null
}
}
The /anthropic endpoints use Anthropic's own format, without code or param; branch on the HTTP status and error.type:
HTTP 404 · /anthropic
{
"type": "error",
"error": {
"type": "not_found_error",
"message": "model: gpt-5.5"
}
}
- Some errors carry extra top-level fields, such as
attempts when every line is busy or retry_after when no line is available for a moment. Ignore any field you do not know.
- When the model provider rejects a request (an over-long context, for example),
message is the provider's own words and code is the provider's value, or invalid_request when it gave none.
HTTP statuses at a glance
| HTTP | When you see it |
| 400 | A bad parameter or an unknown model (such as model_not_found); error.param names the field. |
| 401 | No key, or the key is invalid or disabled (missing_api_key, invalid_api_key). |
| 402 | Not enough balance, or the key has reached its budget (insufficient_quota). |
| 403 | The model is not on this key's model allowlist, or the request comes from outside the key's IP allowlist (permission_denied). |
| 404 | The path does not exist; or the job id is unknown, past its 3-day record, or was submitted by another account (job_not_found). |
| 429 | Over the per-minute limit set on this key or a free model's daily quota, or every line for the model is busy (rate_limit_exceeded). Wait the seconds in Retry-After when it is there, otherwise a few seconds, then retry. |
| 502 | The upstream failed or timed out and switching lines did not help; code says why, e.g. upstream_timeout or upstream_unavailable. Not billed; safe to retry. |
| 503 | The model is under maintenance (model_maintenance), or something failed on our side for a moment (service_unavailable). Not billed; retry later. |
Every error code
The errors the gateway returns, by category. One HTTP status can carry several codes, so go by code.
Authentication and account
| HTTP | code | Cause | What to do |
| 401 | missing_api_key | No Authorization: Bearer header. The /v1 endpoints do not read x-api-key. | Add Authorization: Bearer YOUR_API_KEY. |
| 401 | invalid_api_key | The key does not exist, was deleted or is disabled, or the account is disabled. | Check the key on the API Keys page of the console. |
| 402 | insufficient_quota | The account balance is used up (below the overdraft floor). | Top up and retry. |
| 403 | permission_denied | The request comes from outside the key's IP allowlist. | Call from an allowed IP, or edit the allowlist in the console. |
Limits set on the key
| HTTP | code | Cause | What to do |
| 429 | rate_limit_exceeded | Over the requests-per-minute limit set on this key (counted over the last 60 seconds). No Retry-After header. | Wait a few seconds and retry, or raise the key's limit in the console. |
| 429 | rate_limit_exceeded | A free model's daily quota is used up (per account, per UTC day). message may suggest a paid model to use instead. | Use it again after 00:00 UTC, or switch to a paid model. |
| 402 | insufficient_quota | The key's daily budget (per UTC day) or its total limit is used up. | Use it again the next day, or raise the limit in the console. |
| 403 | permission_denied | The model is not on the key's model allowlist. | Use a model on the allowlist, or edit the allowlist in the console. |
The request itself
| HTTP | code | Cause | What to do |
| 400 | invalid_json / invalid_encoding | The request body is not valid JSON; invalid_encoding when it is not UTF-8. | Check the JSON syntax and the encoding. |
| 400 | invalid_request | A parameter is not valid: the body is not a JSON object, n is not an integer, a reference image or video cannot be read (message says which one and why), /v1/images/edits has no reference image, a video's duration / resolution / ratio / number of references is outside the model's range, and so on. | Fix what message says and submit again. |
| 400 | invalid_messages / invalid_role / invalid_max_tokens | A chat request's messages is missing, empty or malformed, a role is not one of the allowed values, or max_tokens is not a positive integer. param names the field. | Fix the field named in param. |
| 400 | missing_prompt | /v1/images/generations has no prompt. | Add a prompt. |
| 400 | unsupported_resolution | The model does not offer the chosen resolution tier (Grok image generation has no 4K, for example). | Pick another tier; each model's doc lists its tiers. |
| 400 | invalid_callback_url | callback_url is not a public http(s) address, or is longer than 2000 characters. | Use an address that can be reached from the internet. |
| 400 | max_tokens_too_small | A reasoning model spent all of max_tokens thinking and wrote no answer (non-streaming chat). | Raise max_tokens and retry. |
| 404 | not_found | The endpoint path does not exist. | Check the path; the API reference lists every endpoint. |
| 404 | job_not_found | The job id is unknown, past its 3-day record, or was submitted by another account; polling an image job on the video endpoint (or the other way round) gives this error too. | Check the job id and the endpoint. The result link is kept for 60 days and is not affected when the job record expires. |
Model status
| HTTP | code | Cause | What to do |
| 400 | model_not_found | The model id is misspelled, model is missing, or the model cannot be used on this endpoint (a chat model on the image endpoint, for example). | Use a model id from GET /v1/models or the API reference. |
| 400 | model_disabled | The model has been retired. | Switch to another model; where there is a replacement, the old model's pages redirect to it. |
| 400 | model_coming_soon | The model is announced but not open for calls yet. | Wait for the launch announcement. |
| 503 | model_maintenance | The model is under maintenance and takes no requests; no other model answers in its place meanwhile. | Retry when maintenance ends; see Status. |
Capacity and holds
| HTTP | code | Cause | What to do |
| 402 | insufficient_quota | Not enough balance to hold this request. The gateway holds the highest possible cost of the request first, so a burst of concurrent calls can trigger it too. | After a burst, wait a few seconds and retry; otherwise top up. |
| 429 | rate_limit_exceeded | Every line for this model is busy and the request was not sent. Chat endpoints only: image and video jobs are never refused but queue instead. The OpenAI-compatible endpoints add Retry-After. | Wait the seconds in Retry-After, then retry. |
| 503 | service_unavailable | A temporary failure on our side, such as storing a reference file; /v1/responses also answers 503, with Retry-After, when no line is available for a moment. | Retry shortly. |
Upstream errors (chat endpoints)
When a chat endpoint (/v1/chat/completions, /v1/responses, /anthropic/v1/messages) fails, the gateway first retries on other lines and returns one of these errors only when every line has failed. None of them is billed.
| HTTP | code | Cause | What to do |
| 502 | upstream_timeout | A line timed out, including taking too long to send the first token. | Retry; use stream: true for long answers. |
| 502 | upstream_unavailable | A line failed or is unavailable for the moment. | Retry shortly. |
| 502 | upstream_busy | The model provider is rate-limiting. A provider's 429 is never passed on as a 429; you get this 502 instead. | Retry shortly. |
| 502 | upstream_connection_error | The connection to the model provider broke. | Retry. |
| 502 | empty_completion | The model provider returned empty content. This is not a judgement of your content. | Retry. |
| 502 | upstream_error | A line answered with an error sentence instead of an answer. | Retry. |
| 400 | content_filter | The model processed the request but returned nothing, usually because the content tripped the provider's safety policy. | Adjust the prompt and retry; the same prompt is usually refused again. |
| 502 | upstream_prompt_filter | The provider's pre-check stopped the prompt. This is not a judgement by the model. | Retry shortly, or reword slightly. |
| 4xx | — | The model provider rejected the request (an over-long context or an unsupported parameter, for example); the status and message are passed on as the provider sent them, and code is the provider's value. | Change the request as message says. |
The Anthropic endpoints
error.type follows Anthropic's standard values by status: 400 invalid_request_error, 401 authentication_error, 402 billing_error, 403 permission_error, 404 not_found_error, 429 rate_limit_error (overloaded_error when every line is busy), 502 api_error, 503 overloaded_error (api_error while the model is under maintenance).
/anthropic takes Claude models only; any other model answers 404 not_found_error.
max_tokens is required and cannot be 0; otherwise you get a 400.
Failed jobs (images and video)
When an image or video job fails after it was accepted, the job endpoint still answers HTTP 200 with status failed; the reason is in error and error.code is one of the values below. The credits held for the job are back in your balance. With a webhook you receive image.failed / video.failed with the same content.
| code | When you see it | What to do |
content_policy | Video: the prompt, a reference, or the generated picture or sound did not pass content review (param is audio when the sound was flagged). | Rewrite the prompt or replace the reference before you submit again; the same input is refused again. |
moderation_blocked | Image: the model provider's safety system refused the prompt or a reference image (the code may also be content_policy_violation). | Rewrite the prompt or replace the reference; resubmitting it unchanged is usually refused again. |
render_failed | Video: no clip came out this time, usually a one-off. | Submitting the same request again usually works. |
invalid_material | Video: the upstream found a problem with the input. It is a catch-all error and often a one-off. | Resubmit as is first; if it keeps happening, check the references and the prompt length. |
render_timeout | Video: queueing or rendering took too long and the upstream gave the job up. | Submit again. |
wait_timeout | Video: no clip arrived within the longest wait the line allows. | Submit again. |
upstream_unstable | Video: the connection broke during rendering. | Submit again. |
result_fetch_failed | The video was rendered, but downloading or storing it kept failing on our side. | Submit again. |
invalid_request | Video: the line does not accept this combination of parameters (duration, ratio and resolution, for example). | Adjust the parameters as message says and submit again. |
upstream_error | Any other failure: every line was tried, the job waited 15 minutes without starting, the upstream refused without a reason, and so on; message says which. | Submitting again usually works. |
The full async jobs guide →
Errors in the middle of a stream
- Once a streaming request has answered HTTP 200 and started sending, a later error cannot change the status: it either arrives as a data event carrying
error, or the stream simply ends.
- Check that a stream is complete: in the OpenAI format look for a
finish_reason and data: [DONE], in the Anthropic format for message_stop. Without them, treat the stream as cut off.
- A stream that breaks off is billed for the usage produced before it stopped; see Billing and refunds.
Are errors billed?
- A request that returns an error is never billed: either the error comes before any charge, or the hold is returned in full.
- Failed image and video jobs are refunded in full.
- The one exception is a stream that has already started: if it breaks off, you pay for the usage produced before it stopped.
When to retry
For a 4xx, fix the request first (top up on a 402, wait a little on a 429); 5xx errors and timeouts can be retried as they are, with exponential backoff, waiting as long as Retry-After says when it is present.
Retry strategy and sample code →