NezhaGateNezhaGate

One error format for everything, what each code means, and what to do when you see it.

Error format

When a /v1 endpoint fails it answers with a 4xx / 5xx status and JSON like the example below: message is for people, type is the broad class, code is a stable machine-readable identifier and param names the field at fault (null when there is none). Branch on the HTTP status and code, never on the message text: it may be English, Chinese or both, and its wording changes.

HTTP 402 · /v1
{
  "error": {
    "message": "Insufficient balance — please top up",
    "type": "insufficient_quota",
    "code": "insufficient_quota",
    "param": null
  }
}

The /anthropic endpoints use Anthropic's own format, without code or param; branch on the HTTP status and error.type:

HTTP 404 · /anthropic
{
  "type": "error",
  "error": {
    "type": "not_found_error",
    "message": "model: gpt-5.5"
  }
}
  • Some errors carry extra top-level fields, such as attempts when every line is busy or retry_after when no line is available for a moment. Ignore any field you do not know.
  • When the model provider rejects a request (an over-long context, for example), message is the provider's own words and code is the provider's value, or invalid_request when it gave none.

HTTP statuses at a glance

HTTPWhen you see it
400A bad parameter or an unknown model (such as model_not_found); error.param names the field.
401No key, or the key is invalid or disabled (missing_api_key, invalid_api_key).
402Not enough balance, or the key has reached its budget (insufficient_quota).
403The model is not on this key's model allowlist, or the request comes from outside the key's IP allowlist (permission_denied).
404The path does not exist; or the job id is unknown, past its 3-day record, or was submitted by another account (job_not_found).
429Over the per-minute limit set on this key or a free model's daily quota, or every line for the model is busy (rate_limit_exceeded). Wait the seconds in Retry-After when it is there, otherwise a few seconds, then retry.
502The upstream failed or timed out and switching lines did not help; code says why, e.g. upstream_timeout or upstream_unavailable. Not billed; safe to retry.
503The model is under maintenance (model_maintenance), or something failed on our side for a moment (service_unavailable). Not billed; retry later.

Every error code

The errors the gateway returns, by category. One HTTP status can carry several codes, so go by code.

Authentication and account

HTTPcodeCauseWhat to do
401missing_api_keyNo Authorization: Bearer header. The /v1 endpoints do not read x-api-key.Add Authorization: Bearer YOUR_API_KEY.
401invalid_api_keyThe key does not exist, was deleted or is disabled, or the account is disabled.Check the key on the API Keys page of the console.
402insufficient_quotaThe account balance is used up (below the overdraft floor).Top up and retry.
403permission_deniedThe request comes from outside the key's IP allowlist.Call from an allowed IP, or edit the allowlist in the console.

Limits set on the key

HTTPcodeCauseWhat to do
429rate_limit_exceededOver the requests-per-minute limit set on this key (counted over the last 60 seconds). No Retry-After header.Wait a few seconds and retry, or raise the key's limit in the console.
429rate_limit_exceededA free model's daily quota is used up (per account, per UTC day). message may suggest a paid model to use instead.Use it again after 00:00 UTC, or switch to a paid model.
402insufficient_quotaThe key's daily budget (per UTC day) or its total limit is used up.Use it again the next day, or raise the limit in the console.
403permission_deniedThe model is not on the key's model allowlist.Use a model on the allowlist, or edit the allowlist in the console.

The request itself

HTTPcodeCauseWhat to do
400invalid_json / invalid_encodingThe request body is not valid JSON; invalid_encoding when it is not UTF-8.Check the JSON syntax and the encoding.
400invalid_requestA parameter is not valid: the body is not a JSON object, n is not an integer, a reference image or video cannot be read (message says which one and why), /v1/images/edits has no reference image, a video's duration / resolution / ratio / number of references is outside the model's range, and so on.Fix what message says and submit again.
400invalid_messages / invalid_role / invalid_max_tokensA chat request's messages is missing, empty or malformed, a role is not one of the allowed values, or max_tokens is not a positive integer. param names the field.Fix the field named in param.
400missing_prompt/v1/images/generations has no prompt.Add a prompt.
400unsupported_resolutionThe model does not offer the chosen resolution tier (Grok image generation has no 4K, for example).Pick another tier; each model's doc lists its tiers.
400invalid_callback_urlcallback_url is not a public http(s) address, or is longer than 2000 characters.Use an address that can be reached from the internet.
400max_tokens_too_smallA reasoning model spent all of max_tokens thinking and wrote no answer (non-streaming chat).Raise max_tokens and retry.
404not_foundThe endpoint path does not exist.Check the path; the API reference lists every endpoint.
404job_not_foundThe job id is unknown, past its 3-day record, or was submitted by another account; polling an image job on the video endpoint (or the other way round) gives this error too.Check the job id and the endpoint. The result link is kept for 60 days and is not affected when the job record expires.

Model status

HTTPcodeCauseWhat to do
400model_not_foundThe model id is misspelled, model is missing, or the model cannot be used on this endpoint (a chat model on the image endpoint, for example).Use a model id from GET /v1/models or the API reference.
400model_disabledThe model has been retired.Switch to another model; where there is a replacement, the old model's pages redirect to it.
400model_coming_soonThe model is announced but not open for calls yet.Wait for the launch announcement.
503model_maintenanceThe model is under maintenance and takes no requests; no other model answers in its place meanwhile.Retry when maintenance ends; see Status.

Capacity and holds

HTTPcodeCauseWhat to do
402insufficient_quotaNot enough balance to hold this request. The gateway holds the highest possible cost of the request first, so a burst of concurrent calls can trigger it too.After a burst, wait a few seconds and retry; otherwise top up.
429rate_limit_exceededEvery line for this model is busy and the request was not sent. Chat endpoints only: image and video jobs are never refused but queue instead. The OpenAI-compatible endpoints add Retry-After.Wait the seconds in Retry-After, then retry.
503service_unavailableA temporary failure on our side, such as storing a reference file; /v1/responses also answers 503, with Retry-After, when no line is available for a moment.Retry shortly.

Upstream errors (chat endpoints)

When a chat endpoint (/v1/chat/completions, /v1/responses, /anthropic/v1/messages) fails, the gateway first retries on other lines and returns one of these errors only when every line has failed. None of them is billed.

HTTPcodeCauseWhat to do
502upstream_timeoutA line timed out, including taking too long to send the first token.Retry; use stream: true for long answers.
502upstream_unavailableA line failed or is unavailable for the moment.Retry shortly.
502upstream_busyThe model provider is rate-limiting. A provider's 429 is never passed on as a 429; you get this 502 instead.Retry shortly.
502upstream_connection_errorThe connection to the model provider broke.Retry.
502empty_completionThe model provider returned empty content. This is not a judgement of your content.Retry.
502upstream_errorA line answered with an error sentence instead of an answer.Retry.
400content_filterThe model processed the request but returned nothing, usually because the content tripped the provider's safety policy.Adjust the prompt and retry; the same prompt is usually refused again.
502upstream_prompt_filterThe provider's pre-check stopped the prompt. This is not a judgement by the model.Retry shortly, or reword slightly.
4xx—The model provider rejected the request (an over-long context or an unsupported parameter, for example); the status and message are passed on as the provider sent them, and code is the provider's value.Change the request as message says.

The Anthropic endpoints

  • error.type follows Anthropic's standard values by status: 400 invalid_request_error, 401 authentication_error, 402 billing_error, 403 permission_error, 404 not_found_error, 429 rate_limit_error (overloaded_error when every line is busy), 502 api_error, 503 overloaded_error (api_error while the model is under maintenance).
  • /anthropic takes Claude models only; any other model answers 404 not_found_error.
  • max_tokens is required and cannot be 0; otherwise you get a 400.

Failed jobs (images and video)

When an image or video job fails after it was accepted, the job endpoint still answers HTTP 200 with status failed; the reason is in error and error.code is one of the values below. The credits held for the job are back in your balance. With a webhook you receive image.failed / video.failed with the same content.

codeWhen you see itWhat to do
content_policyVideo: the prompt, a reference, or the generated picture or sound did not pass content review (param is audio when the sound was flagged).Rewrite the prompt or replace the reference before you submit again; the same input is refused again.
moderation_blockedImage: the model provider's safety system refused the prompt or a reference image (the code may also be content_policy_violation).Rewrite the prompt or replace the reference; resubmitting it unchanged is usually refused again.
render_failedVideo: no clip came out this time, usually a one-off.Submitting the same request again usually works.
invalid_materialVideo: the upstream found a problem with the input. It is a catch-all error and often a one-off.Resubmit as is first; if it keeps happening, check the references and the prompt length.
render_timeoutVideo: queueing or rendering took too long and the upstream gave the job up.Submit again.
wait_timeoutVideo: no clip arrived within the longest wait the line allows.Submit again.
upstream_unstableVideo: the connection broke during rendering.Submit again.
result_fetch_failedThe video was rendered, but downloading or storing it kept failing on our side.Submit again.
invalid_requestVideo: the line does not accept this combination of parameters (duration, ratio and resolution, for example).Adjust the parameters as message says and submit again.
upstream_errorAny other failure: every line was tried, the job waited 15 minutes without starting, the upstream refused without a reason, and so on; message says which.Submitting again usually works.

The full async jobs guide →

Errors in the middle of a stream

  • Once a streaming request has answered HTTP 200 and started sending, a later error cannot change the status: it either arrives as a data event carrying error, or the stream simply ends.
  • Check that a stream is complete: in the OpenAI format look for a finish_reason and data: [DONE], in the Anthropic format for message_stop. Without them, treat the stream as cut off.
  • A stream that breaks off is billed for the usage produced before it stopped; see Billing and refunds.

Are errors billed?

  • A request that returns an error is never billed: either the error comes before any charge, or the hold is returned in full.
  • Failed image and video jobs are refunded in full.
  • The one exception is a stream that has already started: if it breaks off, you pay for the usage produced before it stopped.

When to retry

For a 4xx, fix the request first (top up on a 402, wait a little on a 429); 5xx errors and timeouts can be retried as they are, with exponential backoff, waiting as long as Retry-After says when it is present.

Retry strategy and sample code →