# Rate limits and retries

Account and per-key limits, what you get when every line is busy, and how to retry and set timeouts.

> https://nezhagate.com/en/docs/guide/rate-limits

## No default rate limit

NezhaGate sets no default request-rate limit per account, and capacity grows with your usage. To cap usage, give each key its own limits.

## Per-key limits

Set these per key on the API Keys page of the console; empty or 0 means no limit:

| Setting | What it does | When exceeded |
| --- | --- | --- |
| **Requests per minute** | The most requests this key may send in a minute. | `429` `rate_limit_exceeded` |
| **Daily budget** | The most this key may spend per day (UTC), in US dollars. | `402` `insufficient_quota` |
| **Total limit** | The most this key may spend in total, in US dollars; good for automation that must not run away. | `402` `insufficient_quota` |
| **Model allowlist** | This key may call only the models listed. | `403` `permission_denied` |
| **IP allowlist** | Only requests from the listed IPs or ranges are accepted. | `403` `permission_denied` |

## When every line is busy

- Each model is served by several lines and the gateway spreads requests across them. If every line for a model is at capacity at once you get a 429; the OpenAI-compatible endpoints add a `Retry-After` header with the seconds to wait.

- Image and video jobs are not refused when busy: they wait in a queue (`status` is `queued`, with the position and expected wait) and start on their own.

- Free promotional models also have a daily call quota; past it you get 429 until it resets at 00:00 UTC.

## When to retry

| HTTP | What to do |
| --- | --- |
| **429** | Wait the seconds in the `Retry-After` header, then retry; without the header, back off 1, 2, then 4 seconds. |
| **500 / 502 / 503 / 504** | Retry with exponential backoff (1, 2, 4, 8 seconds plus jitter); 3–5 tries are usually enough. These failures are not billed. |
| **Other 4xx** | Retrying unchanged will not help: fix the request according to `error.code` (parameters, model, key, balance) first. |
| **Network error / timeout** | Chat can simply be retried. For images and video, check the call log for the job first, then decide whether to submit again. |

```
import random
import time
import requests

def post_with_retry(url, headers, payload, tries=5):
    for attempt in range(tries):
        r = requests.post(url, headers=headers, json=payload, timeout=600)
        if r.status_code == 429:
            time.sleep(float(r.headers.get("Retry-After", 2 ** attempt)))
            continue
        if r.status_code >= 500:
            time.sleep(2 ** attempt + random.random())     # 1, 2, 4, 8 s ... plus jitter
            continue
        return r                                           # 2xx, or a 4xx to fix in the request
    return r
```

## Timeouts

- A single non-streaming request can run for about 10 minutes before the edge gateway cuts it (HTTP 524). For long answers and reasoning models, use `stream: true`.

- Set your client timeout to at least 10 minutes, or your side gives up before a long answer arrives.

- Image and video are async jobs: the submit call returns quickly, so this limit does not apply to them.

## Need more capacity

For sustained high concurrency, tell us your volume ahead of time in the [Telegram group](https://t.me/+-erBoH9-AYY4MTM1) or through support, and we will add capacity for it.
