NezhaGateNezhaGate

Every endpoint uses the same API key and the same balance. Chat speaks the OpenAI and Anthropic formats; images and video run as async jobs.

Base URLs

Pick the one that matches your SDK. Both share one key and one balance.

ProtocolBase URLUse it for
OpenAI-compatiblehttps://nezhagate.com/v1Chat, Responses, images, video, model list, balance
Native Anthropichttps://nezhagate.com/anthropicThe Messages API for Claude models (Claude Code, the Anthropic SDK)

Authentication

Send your API key in a header on every request. Create keys on the API Keys page of the console; a key is shown only once, when you create it.

Header
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json

The Anthropic endpoints take an x-api-key header and also accept Authorization: Bearer.

Manage API keys →

Endpoints

All paths are under https://nezhagate.com. Synchronous endpoints return the result; image and video endpoints return a job id.

MethodPathWhat it does
POST/v1/chat/completionsChat with any chat model. Streams with stream: true.
POST/v1/responsesChat in the OpenAI Responses format (selected models).
POST/anthropic/v1/messagesThe native Claude Messages API, with streaming, tool use and prompt caching.
POST/anthropic/v1/messages/count_tokensCount the tokens of a Messages request.
GET/anthropic/v1/modelsList Claude models (Anthropic format).
POST/v1/images/generationsText-to-image. Returns a job id at once (HTTP 202).
POST/v1/images/editsImage-to-image. Passing image to the call above does the same.
GET/v1/images/jobs/{id}Get an image job: status and result.
POST/v1/videos/generationsGenerate a video. Returns a job id at once (HTTP 202).
GET/v1/videos/jobs/{id}Get a video job: status and result.
GET/v1/modelsList the models on sale (OpenAI format).
GET/v1/usageBalance, total and today's spend, and usage per model.
GET/v1/dashboard/billing/credit_grantsOpenAI-style balance check; the balance is total_available.

Async jobs (images and video)

Image and video calls return HTTP 202 with a job id at once; the work runs in the background. Poll with the id, or have the result pushed to you by webhook.

StatusMeaning
queuedWaiting in line. Carries queue_position (jobs ahead of yours) and eta_seconds (expected wait).
processingRendering.
succeededDone. The result is in data[0].url, the billing detail in usage.
failedFailed. The reason is in error; the credits held for the job are refunded in full.
curl · Submit a job
curl https://nezhagate.com/v1/images/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "nano-banana-2", "prompt": "a lighthouse at dawn, watercolor", "size": "16:9"}'

# HTTP 202
{"id": "img_3f9a...c2", "object": "image.generation.job", "status": "queued", "model": "nano-banana-2"}
curl · Poll the job
curl https://nezhagate.com/v1/images/jobs/img_3f9a...c2 -H "Authorization: Bearer YOUR_API_KEY"
Response when done
{
  "id": "img_3f9a...c2",
  "object": "image.generation.job",
  "status": "succeeded",
  "created": 1791199400,
  "model": "nano-banana-2",
  "data": [{"url": "https://img.nezhagate.com/i/9f86d081a8....png"}],
  "usage": {"images": 1, "resolution": "1K", "model": "nano-banana-2"}
}
  • Poll images every 2–3 seconds and video every 5–10 seconds. Polling is free.
  • data[0].url is a link on our media host, kept for 60 days and then deleted; download anything you need to keep.
  • Add callback_url when you submit, and the same result the poll returns is pushed to that URL when the job ends. Webhook docs →

List models

Returns every model ID on sale, in the OpenAI format. Retired models are not listed; a coming-soon model is listed but answers 400 model_coming_soon until it opens.

curl · GET /v1/models
curl https://nezhagate.com/v1/models -H "Authorization: Bearer YOUR_API_KEY"
200 · JSON
{"object": "list", "data": [{"id": "gpt-5.5", "object": "model", "owned_by": "..."}, {"id": "claude-sonnet-5", "object": "model", "owned_by": "..."}]}

Balance & usage

Any API key can read the account's balance and spend. No console login needed.

curl · GET /v1/usage
curl https://nezhagate.com/v1/usage -H "Authorization: Bearer YOUR_API_KEY"
200 · JSON
{
  "object": "usage",
  "balance": {"usd": 12.5, "credits": 2500},
  "total": {"cost_usd": 37.5, "requests": 1840, "billed_requests": 1822, "failed_requests": 18},
  "today": {"cost_usd": 1.2, "requests": 64, "billed_requests": 63, "failed_requests": 1},
  "by_model": [
    {"model": "gpt-5.5", "cost_usd": 20.1, "requests": 900, "billed_requests": 896, "failed_requests": 4,
     "prompt_tokens": 1520000, "completion_tokens": 410000, "image_count": 0}
  ]
}

balance.usd is the balance in US dollars and balance.credits the same in credits (1 USD = 200 credits).

For tools that already know how to check an OpenAI balance, use the OpenAI-style endpoint; the balance is total_available.

curl · GET /v1/dashboard/billing/credit_grants
curl https://nezhagate.com/v1/dashboard/billing/credit_grants -H "Authorization: Bearer YOUR_API_KEY"

Error format

Every error has the same shape. code is a stable, machine-readable identifier: branch on code, never on the message text.

400 · JSON
{
  "error": {
    "message": "Model not enabled: gpt-9",
    "type": "invalid_request_error",
    "code": "model_not_found",
    "param": "model"
  }
}
HTTPWhen you see it
400A bad parameter or an unknown model (such as model_not_found); error.param names the field.
401No key, or the key is invalid or disabled (missing_api_key, invalid_api_key).
402Not enough balance, or the key has reached its budget (insufficient_quota).
403The model is not on this key's model allowlist.
404No such job, or it belongs to another key (job_not_found).
429Rate limited, or every line for this model is busy. Wait the seconds in the Retry-After header, then retry.
502The upstream failed or timed out (upstream_error). Not billed; safe to retry.
503The model is under maintenance (model_maintenance) and takes no orders until it is back.

Full list of error codes →

Rate limits & retries

  • Accounts have no default request-rate limit. In the console you can give each key its own requests-per-minute cap, daily budget, total spend limit, model allowlist and IP allowlist.
  • On a 429, wait as long as the Retry-After header says, then retry. For 502, 503 and timeouts, retry with exponential backoff (for example 1, 2, then 4 seconds).
  • Failed requests are never billed. If a stream breaks off midway, you pay only for what was returned.
  • For sustained high concurrency, tell us ahead of time and we will add capacity for your volume.

SDKs & examples

No special SDK needed: the official OpenAI and Anthropic SDKs work once you change base_url.

Python · openai
from openai import OpenAI

client = OpenAI(base_url="https://nezhagate.com/v1", api_key="YOUR_API_KEY")
resp = client.chat.completions.create(
    model="gpt-5.5",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)
for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="")

More runnable examples (Python, Node.js, curl, including the async image and video flow): github.com/gaoorange/nezhagate-api-examples

Docs for each model

Every model has its own page: allowed parameter values, example code, the response format and how it is billed.

Chat models

Image models

Video models