NezhaGateNezhaGate
POST https://nezhagate.com/v1/chat/completions

OpenAI-compatible chat completions backed by Alibaba Cloud Qwen3.8 Flash, the light, fast member of the Qwen3.8 family with a low unit price for batch jobs and high concurrency. It always thinks; the reasoning returns in message.reasoning_content and reasoning tokens are billed at the output rate within usage.completion_tokens. Tool calling, JSON output, streaming and image input (public links or base64) are supported; cache_control {"type": "ephemeral"} is accepted and the upstream may serve a repeated long prefix from cache, but whether it hits is decided upstream and billing follows the returned usage (a reported cache hit settles at the cache rate). Set model to qwen3.8-flash; served on /v1/chat/completions only.

Try in Playground →

Authentication

Header
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json

Create an API Key in the console to start.

Request body

ParameterTypeRequiredDescription
model string Yes Model ID, here qwen3.8-flash.
messages array Yes Array of messages; each has role (system/user/assistant) and content. content may be a string, or an array of {type:text} and {type:image_url} parts for image understanding (multimodal/vision).
stream boolean No Stream the response as SSE. Default false.
temperature number No Sampling temperature, 0–2.
max_tokens integer No Maximum number of tokens to generate.
web_search boolean No Set true to enable web search: the gateway augments the prompt with live results (citing sources) before the model answers. Can also be triggered via a tools entry {"type":"web_search"}.

Request example

cURL
curl https://nezhagate.com/v1/chat/completions -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d '{"model": "qwen3.8-flash", "messages": [{"role": "user", "content": "Hello"}], "stream": false}'

Response

200 · JSON
{
  "id": "chatcmpl_xxx",
  "object": "chat.completion",
  "model": "qwen3.8-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": "The user is greeting me, so a short friendly reply fits...",
        "content": "Hello! How can I help you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 9,
    "completion_tokens": 42,
    "total_tokens": 51,
    "completion_tokens_details": {"reasoning_tokens": 31}
  }
}

Image input (Vision)

Put an image in the message content array and the model will analyze it (visual Q&A, reading text / OCR, …). image_url accepts a public image link or an inline base64 data URL (data:image/png;base64,...). Available on multimodal models (gpt-5.5, gemini series, …).

curl · Image input (Vision)
curl https://nezhagate.com/v1/chat/completions -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d '{"model": "qwen3.8-flash", "messages": [{"role": "user", "content": [{"type": "text", "text": "What is in this picture?"}, {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}]}]}'

Error codes

Every error body carries error.message / error.type / error.code / error.param - branch on code; the full list is in the integration guide.

HTTPcodeDescription
401invalid_api_keyAPI Key missing or invalid
402insufficient_quotaInsufficient balance or Key over quota
400invalid_requestUnsupported model or parameter
429rate_limit_exceededUpstream rate limit
502upstream_errorAll upstream routes failed