NezhaGateNezhaGate

Account and per-key limits, what you get when every line is busy, and how to retry and set timeouts.

No default rate limit

NezhaGate sets no default request-rate limit per account, and capacity grows with your usage. To cap usage, give each key its own limits.

Per-key limits

Set these per key on the API Keys page of the console; empty or 0 means no limit:

SettingWhat it doesWhen exceeded
Requests per minuteThe most requests this key may send in a minute.429 rate_limit_exceeded
Daily budgetThe most this key may spend per day (UTC), in US dollars.402 insufficient_quota
Total limitThe most this key may spend in total, in US dollars; good for automation that must not run away.402 insufficient_quota
Model allowlistThis key may call only the models listed.403 permission_denied
IP allowlistOnly requests from the listed IPs or ranges are accepted.403 permission_denied

When every line is busy

  • Each model is served by several lines and the gateway spreads requests across them. If every line for a model is at capacity at once you get a 429; the OpenAI-compatible endpoints add a Retry-After header with the seconds to wait.
  • Image and video jobs are not refused when busy: they wait in a queue (status is queued, with the position and expected wait) and start on their own.
  • Free promotional models also have a daily call quota; past it you get 429 until it resets at 00:00 UTC.

When to retry

HTTPWhat to do
429Wait the seconds in the Retry-After header, then retry; without the header, back off 1, 2, then 4 seconds.
500 / 502 / 503 / 504Retry with exponential backoff (1, 2, 4, 8 seconds plus jitter); 3–5 tries are usually enough. These failures are not billed.
Other 4xxRetrying unchanged will not help: fix the request according to error.code (parameters, model, key, balance) first.
Network error / timeoutChat can simply be retried. For images and video, check the call log for the job first, then decide whether to submit again.
Python · retrying 429 and 5xx
import random
import time
import requests


def post_with_retry(url, headers, payload, tries=5):
    for attempt in range(tries):
        r = requests.post(url, headers=headers, json=payload, timeout=600)
        if r.status_code == 429:
            time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
            continue
        if r.status_code >= 500:
            time.sleep(2 ** attempt + random.random())     # 1, 2, 4, 8 s ... plus jitter
            continue
        return r                                           # 2xx, or a 4xx to fix in the request
    return r

Timeouts

  • A single non-streaming request can run for about 10 minutes before the edge gateway cuts it (HTTP 524). For long answers and reasoning models, use stream: true.
  • Set your client timeout to at least 10 minutes, or your side gives up before a long answer arrives.
  • Image and video are async jobs: the submit call returns quickly, so this limit does not apply to them.

Need more capacity

For sustained high concurrency, tell us your volume ahead of time in the Telegram group or through support, and we will add capacity for it.