Rate limits and retries
Account and per-key limits, what you get when every line is busy, and how to retry and set timeouts.
No default rate limit
NezhaGate sets no default request-rate limit per account, and capacity grows with your usage. To cap usage, give each key its own limits.
Per-key limits
Set these per key on the API Keys page of the console; empty or 0 means no limit:
| Setting | What it does | When exceeded |
|---|---|---|
| Requests per minute | The most requests this key may send in a minute. | 429 rate_limit_exceeded |
| Daily budget | The most this key may spend per day (UTC), in US dollars. | 402 insufficient_quota |
| Total limit | The most this key may spend in total, in US dollars; good for automation that must not run away. | 402 insufficient_quota |
| Model allowlist | This key may call only the models listed. | 403 permission_denied |
| IP allowlist | Only requests from the listed IPs or ranges are accepted. | 403 permission_denied |
When every line is busy
- Each model is served by several lines and the gateway spreads requests across them. If every line for a model is at capacity at once you get a 429; the OpenAI-compatible endpoints add a
Retry-Afterheader with the seconds to wait. - Image and video jobs are not refused when busy: they wait in a queue (
statusisqueued, with the position and expected wait) and start on their own. - Free promotional models also have a daily call quota; past it you get 429 until it resets at 00:00 UTC.
When to retry
| HTTP | What to do |
|---|---|
| 429 | Wait the seconds in the Retry-After header, then retry; without the header, back off 1, 2, then 4 seconds. |
| 500 / 502 / 503 / 504 | Retry with exponential backoff (1, 2, 4, 8 seconds plus jitter); 3–5 tries are usually enough. These failures are not billed. |
| Other 4xx | Retrying unchanged will not help: fix the request according to error.code (parameters, model, key, balance) first. |
| Network error / timeout | Chat can simply be retried. For images and video, check the call log for the job first, then decide whether to submit again. |
import random
import time
import requests
def post_with_retry(url, headers, payload, tries=5):
for attempt in range(tries):
r = requests.post(url, headers=headers, json=payload, timeout=600)
if r.status_code == 429:
time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
continue
if r.status_code >= 500:
time.sleep(2 ** attempt + random.random()) # 1, 2, 4, 8 s ... plus jitter
continue
return r # 2xx, or a 4xx to fix in the request
return rTimeouts
- A single non-streaming request can run for about 10 minutes before the edge gateway cuts it (HTTP 524). For long answers and reasoning models, use
stream: true. - Set your client timeout to at least 10 minutes, or your side gives up before a long answer arrives.
- Image and video are async jobs: the submit call returns quickly, so this limit does not apply to them.
Need more capacity
For sustained high concurrency, tell us your volume ahead of time in the Telegram group or through support, and we will add capacity for it.