API reference
Every endpoint uses the same API key and the same balance. Chat speaks the OpenAI and Anthropic formats; images and video run as async jobs.
Base URLs
Pick the one that matches your SDK. Both share one key and one balance.
| Protocol | Base URL | Use it for |
|---|---|---|
| OpenAI-compatible | https://nezhagate.com/v1 | Chat, Responses, images, video, model list, balance |
| Native Anthropic | https://nezhagate.com/anthropic | The Messages API for Claude models (Claude Code, the Anthropic SDK) |
Authentication
Send your API key in a header on every request. Create keys on the API Keys page of the console; a key is shown only once, when you create it.
Authorization: Bearer YOUR_API_KEY Content-Type: application/json
The Anthropic endpoints take an x-api-key header and also accept Authorization: Bearer.
Endpoints
All paths are under https://nezhagate.com. Synchronous endpoints return the result; image and video endpoints return a job id.
| Method | Path | What it does |
|---|---|---|
| POST | /v1/chat/completions | Chat with any chat model. Streams with stream: true. |
| POST | /v1/responses | Chat in the OpenAI Responses format (selected models). |
| POST | /anthropic/v1/messages | The native Claude Messages API, with streaming, tool use and prompt caching. |
| POST | /anthropic/v1/messages/count_tokens | Count the tokens of a Messages request. |
| GET | /anthropic/v1/models | List Claude models (Anthropic format). |
| POST | /v1/images/generations | Text-to-image. Returns a job id at once (HTTP 202). |
| POST | /v1/images/edits | Image-to-image. Passing image to the call above does the same. |
| GET | /v1/images/jobs/{id} | Get an image job: status and result. |
| POST | /v1/videos/generations | Generate a video. Returns a job id at once (HTTP 202). |
| GET | /v1/videos/jobs/{id} | Get a video job: status and result. |
| GET | /v1/models | List the models on sale (OpenAI format). |
| GET | /v1/usage | Balance, total and today's spend, and usage per model. |
| GET | /v1/dashboard/billing/credit_grants | OpenAI-style balance check; the balance is total_available. |
Async jobs (images and video)
Image and video calls return HTTP 202 with a job id at once; the work runs in the background. Poll with the id, or have the result pushed to you by webhook.
| Status | Meaning |
|---|---|
queued | Waiting in line. Carries queue_position (jobs ahead of yours) and eta_seconds (expected wait). |
processing | Rendering. |
succeeded | Done. The result is in data[0].url, the billing detail in usage. |
failed | Failed. The reason is in error; the credits held for the job are refunded in full. |
curl https://nezhagate.com/v1/images/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "nano-banana-2", "prompt": "a lighthouse at dawn, watercolor", "size": "16:9"}'
# HTTP 202
{"id": "img_3f9a...c2", "object": "image.generation.job", "status": "queued", "model": "nano-banana-2"}curl https://nezhagate.com/v1/images/jobs/img_3f9a...c2 -H "Authorization: Bearer YOUR_API_KEY"
{
"id": "img_3f9a...c2",
"object": "image.generation.job",
"status": "succeeded",
"created": 1791199400,
"model": "nano-banana-2",
"data": [{"url": "https://img.nezhagate.com/i/9f86d081a8....png"}],
"usage": {"images": 1, "resolution": "1K", "model": "nano-banana-2"}
}- Poll images every 2–3 seconds and video every 5–10 seconds. Polling is free.
data[0].urlis a link on our media host, kept for 60 days and then deleted; download anything you need to keep.- Add
callback_urlwhen you submit, and the same result the poll returns is pushed to that URL when the job ends. Webhook docs →
List models
Returns every model ID on sale, in the OpenAI format. Retired models are not listed; a coming-soon model is listed but answers 400 model_coming_soon until it opens.
curl https://nezhagate.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"
{"object": "list", "data": [{"id": "gpt-5.5", "object": "model", "owned_by": "..."}, {"id": "claude-sonnet-5", "object": "model", "owned_by": "..."}]}Balance & usage
Any API key can read the account's balance and spend. No console login needed.
curl https://nezhagate.com/v1/usage -H "Authorization: Bearer YOUR_API_KEY"
{
"object": "usage",
"balance": {"usd": 12.5, "credits": 2500},
"total": {"cost_usd": 37.5, "requests": 1840, "billed_requests": 1822, "failed_requests": 18},
"today": {"cost_usd": 1.2, "requests": 64, "billed_requests": 63, "failed_requests": 1},
"by_model": [
{"model": "gpt-5.5", "cost_usd": 20.1, "requests": 900, "billed_requests": 896, "failed_requests": 4,
"prompt_tokens": 1520000, "completion_tokens": 410000, "image_count": 0}
]
}balance.usd is the balance in US dollars and balance.credits the same in credits (1 USD = 200 credits).
For tools that already know how to check an OpenAI balance, use the OpenAI-style endpoint; the balance is total_available.
curl https://nezhagate.com/v1/dashboard/billing/credit_grants -H "Authorization: Bearer YOUR_API_KEY"
Error format
Every error has the same shape. code is a stable, machine-readable identifier: branch on code, never on the message text.
{
"error": {
"message": "Model not enabled: gpt-9",
"type": "invalid_request_error",
"code": "model_not_found",
"param": "model"
}
}| HTTP | When you see it |
|---|---|
| 400 | A bad parameter or an unknown model (such as model_not_found); error.param names the field. |
| 401 | No key, or the key is invalid or disabled (missing_api_key, invalid_api_key). |
| 402 | Not enough balance, or the key has reached its budget (insufficient_quota). |
| 403 | The model is not on this key's model allowlist. |
| 404 | No such job, or it belongs to another key (job_not_found). |
| 429 | Rate limited, or every line for this model is busy. Wait the seconds in the Retry-After header, then retry. |
| 502 | The upstream failed or timed out (upstream_error). Not billed; safe to retry. |
| 503 | The model is under maintenance (model_maintenance) and takes no orders until it is back. |
Rate limits & retries
- Accounts have no default request-rate limit. In the console you can give each key its own requests-per-minute cap, daily budget, total spend limit, model allowlist and IP allowlist.
- On a 429, wait as long as the
Retry-Afterheader says, then retry. For 502, 503 and timeouts, retry with exponential backoff (for example 1, 2, then 4 seconds). - Failed requests are never billed. If a stream breaks off midway, you pay only for what was returned.
- For sustained high concurrency, tell us ahead of time and we will add capacity for your volume.
SDKs & examples
No special SDK needed: the official OpenAI and Anthropic SDKs work once you change base_url.
from openai import OpenAI
client = OpenAI(base_url="https://nezhagate.com/v1", api_key="YOUR_API_KEY")
resp = client.chat.completions.create(
model="gpt-5.5",
messages=[{"role": "user", "content": "Hello"}],
stream=True,
)
for chunk in resp:
print(chunk.choices[0].delta.content or "", end="")More runnable examples (Python, Node.js, curl, including the async image and video flow): github.com/gaoorange/nezhagate-api-examples
Docs for each model
Every model has its own page: allowed parameter values, example code, the response format and how it is billed.
Chat models
- GPT-5.6 Sol
- GPT-5.6 Terra
- GPT-5.6 Luna
- GPT-5.5
- GPT-6 Astra
- GPT-6.1 Sol
- GPT-6 Sol
- GPT-6 Luna
- Claude Sonnet 4.6
- Claude Opus 5
- Claude Fable 5
- Claude Sonnet 5
- Claude Opus 5.5
- Claude Sonnet 5.5
- Gemini 3.1 Pro
- Gemini 3.8 Flash
- Gemini 3.7 Flash
- Gemini 3.6 Flash
- Gemini 3.6 Flash High
- Gemini 3.6 Flash Low
- Gemini 3.6 Flash Tiered
- Gemini 3 Flash
- Gemini 2.5 Flash
- DeepSeek V4.1 Flash
- DeepSeek V4 Flash 0731
- GLM-5.3
- GLM-5.3 Flash
- Kimi K3
- Qwen3.7 Max
- Qwen3.8 Max
- Qwen3.8 Max 0902
- Qwen3.8 Flash
- Doubao Seed 2.1 Pro
- Doubao Seed 2.1 Turbo
- Grok 4.7
Image models
- GPT Image 2
- GPT Image 2.5 Flare
- GPT Image 2.5 Sunburst
- Nano Banana 2
- Nano Banana Pro
- Grok Imagine Image
- Grok Imagine Image Quality