# Qwen3.8 Flash — /chat/completions

OpenAI-compatible chat completions backed by Alibaba Cloud Qwen3.8 Flash, the light, fast member of the Qwen3.8 family with a low unit price for batch jobs and high concurrency. It always thinks; the reasoning returns in message.reasoning_content and reasoning tokens are billed at the output rate within usage.completion_tokens. Tool calling, JSON output, streaming and image input (public links or base64) are supported; cache_control {"type": "ephemeral"} is accepted and the upstream may serve a repeated long prefix from cache, but whether it hits is decided upstream and billing follows the returned usage (a reported cache hit settles at the cache rate). Set model to qwen3.8-flash; served on /v1/chat/completions only.

**Endpoint:** `POST https://nezhagate.com/v1/chat/completions`

## Authentication
```
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json
```

## Request body
| Parameter | Type | Required | Description |
| --- | --- | --- | --- |
| `model` | string | Yes | Model ID, here qwen3.8-flash. |
| `messages` | array | Yes | Array of messages; each has role (system/user/assistant) and content. content may be a string, or an array of {type:text} and {type:image_url} parts for image understanding (multimodal/vision). |
| `stream` | boolean | No | Stream the response as SSE. Default false. |
| `temperature` | number | No | Sampling temperature, 0–2. |
| `max_tokens` | integer | No | Maximum number of tokens to generate. |
| `web_search` | boolean | No | Set true to enable web search: the gateway augments the prompt with live results (citing sources) before the model answers. Can also be triggered via a tools entry {"type":"web_search"}. |

## Request example
```bash
curl https://nezhagate.com/v1/chat/completions -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d '{"model": "qwen3.8-flash", "messages": [{"role": "user", "content": "Hello"}], "stream": false}'
```

## Response
```json
{
  "id": "chatcmpl_xxx",
  "object": "chat.completion",
  "model": "qwen3.8-flash",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "reasoning_content": "The user is greeting me, so a short friendly reply fits...",
        "content": "Hello! How can I help you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 9,
    "completion_tokens": 42,
    "total_tokens": 51,
    "completion_tokens_details": {"reasoning_tokens": 31}
  }
}
```

## Image input (Vision)
Put an image in the message content array and the model will analyze it (visual Q&A, reading text / OCR, …). image_url accepts a public image link or an inline base64 data URL (data:image/png;base64,...). Available on multimodal models (gpt-5.5, gemini series, …).

```bash
curl https://nezhagate.com/v1/chat/completions -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d '{"model": "qwen3.8-flash", "messages": [{"role": "user", "content": [{"type": "text", "text": "What is in this image?"}, {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}]}]}'
```