NezhaGateNezhaGate

How chat, images and video are billed, how holds settle, when nothing is charged, and how to check spend, top up and get refunds.

Credits

Your balance is shown in credits: 1 US dollar = 200 credits. Prices on the pricing page and in the docs are in US dollars and are converted to credits at that rate when you are charged. /v1/usage returns both the dollar balance and the credits.

How each kind of model is charged

KindBilled byNotes
ChatTokens: input and output priced separately (USD per million tokens)Cache hits are billed at the lower cache-read price; reasoning tokens at the output price.
ImagePer image, priced by resolution tier (1K / 2K / 4K)You pay for the images actually delivered; failed ones are free.
VideoMostly per second (rate × duration); a few models per clipWhether a model bills per second or per clip is stated on the pricing page and in its doc.

Prices for every model →

Chat in detail

  • usage.prompt_tokens is the input and usage.completion_tokens the output; cache hits are billed separately at the cache-read price.
  • On reasoning models the thinking tokens count as output and are billed at the output price whether or not you read the thinking; most models report how many in usage.completion_tokens_details.reasoning_tokens.
  • If a stream breaks off midway, you pay for the usage produced before it stopped, not for a full answer.
  • Claude prompt caching also has a cache-write price; each model's cache-read and cache-write prices are in its doc and on the pricing page.

Holds and settlement

  • Before a request runs, an estimate is held from your balance. Without enough balance you get 402 insufficient_quota and the model is never called.
  • Chat settles on actual usage and returns the rest of the hold; images settle by the number of pictures delivered; video settles the duration or clip you ordered, the same as the hold.
  • The more requests run at once, the more is held. On a tight balance a request can get 402 because of holds; it clears as earlier requests settle.

Failures are free

No failed request is ever charged, whether the upstream failed, timed out, refused the content or the gateway erred. A failed async job returns its full hold automatically, with the reason in the call log.

When another model answers

If some chat models are briefly unavailable, the gateway may have another model answer so the request does not fail. You then pay the lower of the two models' prices, and the call log says which model answered and which price applied.

Checking what you spent

  • Console → Call logs: every call from the last 60 days with model, time, usage and charge; failed calls show the reason.
  • Console → Usage: totals by day, model and key, with CSV export.
  • In code: GET /v1/usage returns the balance, total and today's spend and usage per model; see the API reference.

Top-ups and refunds

Top up on the Billing page of the console: from ¥10 (260 credits) with WeChat Pay, from $5 with USDT or a card. Credits do not expire. Sign-up and first top-up bonuses are explained in the FAQ.

Within 7 days of a top-up, unused purchased credits can be refunded through support, minus a $0.40 fee per refund; bonus credits are not refundable. USDT payments can be refunded too.