OpenAI-compatible chat completions backed by Alibaba Cloud Qwen3.8 Flash, the light, fast member of the Qwen3.8 family with a low unit price for batch jobs and high concurrency. It always thinks; the reasoning returns in message.reasoning_content and reasoning tokens are billed at the output rate within usage.completion_tokens. Tool calling, JSON output, streaming and image input (public links or base64) are supported; cache_control {"type": "ephemeral"} is accepted and the upstream may serve a repeated long prefix from cache, but whether it hits is decided upstream and billing follows the returned usage (a reported cache hit settles at the cache rate). Set model to qwen3.8-flash; served on /v1/chat/completions only.
Try in Playground →Authentication
Authorization: Bearer YOUR_API_KEY Content-Type: application/json
Create an API Key in the console to start.
Request body
| Parameter | Type | Required | Description |
|---|---|---|---|
| model | string | Yes | Model ID, here qwen3.8-flash. |
| messages | array | Yes | Array of messages; each has role (system/user/assistant) and content. content may be a string, or an array of {type:text} and {type:image_url} parts for image understanding (multimodal/vision). |
| stream | boolean | No | Stream the response as SSE. Default false. |
| temperature | number | No | Sampling temperature, 0–2. |
| max_tokens | integer | No | Maximum number of tokens to generate. |
| web_search | boolean | No | Set true to enable web search: the gateway augments the prompt with live results (citing sources) before the model answers. Can also be triggered via a tools entry {"type":"web_search"}. |
Request example
curl https://nezhagate.com/v1/chat/completions -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d '{"model": "qwen3.8-flash", "messages": [{"role": "user", "content": "Hello"}], "stream": false}'Response
{
"id": "chatcmpl_xxx",
"object": "chat.completion",
"model": "qwen3.8-flash",
"choices": [
{
"index": 0,
"message": {
"role": "assistant",
"reasoning_content": "The user is greeting me, so a short friendly reply fits...",
"content": "Hello! How can I help you today?"
},
"finish_reason": "stop"
}
],
"usage": {
"prompt_tokens": 9,
"completion_tokens": 42,
"total_tokens": 51,
"completion_tokens_details": {"reasoning_tokens": 31}
}
}Image input (Vision)
Put an image in the message content array and the model will analyze it (visual Q&A, reading text / OCR, …). image_url accepts a public image link or an inline base64 data URL (data:image/png;base64,...). Available on multimodal models (gpt-5.5, gemini series, …).
curl https://nezhagate.com/v1/chat/completions -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d '{"model": "qwen3.8-flash", "messages": [{"role": "user", "content": [{"type": "text", "text": "What is in this picture?"}, {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}]}]}'Error codes
Every error body carries error.message / error.type / error.code / error.param - branch on code; the full list is in the integration guide.
| HTTP | code | Description |
|---|---|---|
| 401 | invalid_api_key | API Key missing or invalid |
| 402 | insufficient_quota | Insufficient balance or Key over quota |
| 400 | invalid_request | Unsupported model or parameter |
| 429 | rate_limit_exceeded | Upstream rate limit |
| 502 | upstream_error | All upstream routes failed |