NezhaGateNezhaGate
Chat

Qwen3.8 Flash

⧉
qwen3.8-flash
Model Type Max0902Flash

Qwen3.8 Flash is the light, fast member of the Qwen3.8 family, with a low unit price and high throughput for batch jobs and high concurrency; short prompts came back in 2.5-6 seconds in our tests. It always thinks before it answers (the reasoning comes back in reasoning_content) and supports tool calling, JSON output, streaming and image input (public image links or base64). cache_control on a text block is accepted, and the upstream may serve a repeated long prefix from cache; whether it hits is decided upstream, and billing follows the returned usage (a reported cache hit settles at the cache rate). Fully OpenAI-compatible: set model to qwen3.8-flash. Pay-as-you-go, failed calls never billed.

ChatLightLow costImage inputOpenAI compatible

Live Test · Playground

Try out Qwen3.8 Flash right here (available after login).

Input

Advanced
Web search
Memory · multi-turn

Conversation

Start a conversationType a message below to begin

About Qwen3.8 Flash

Qwen3.8 Flash is the light, fast member of Alibaba Cloud's Qwen3.8 family, served on NezhaGate through the OpenAI-compatible API. It has a low unit price and high throughput for batch jobs and high concurrency; short prompts came back in 2.5-6 seconds in our tests. It always thinks before it answers, with the reasoning in message.reasoning_content. Tool calling, JSON output, streaming and image input (public links or base64) are supported; cache_control is accepted too, and the upstream may serve a repeated long prefix from cache: whether it hits is decided upstream, and billing follows the returned usage (a reported cache hit settles at the cache rate). Set model to qwen3.8-flash; pay-as-you-go, failed calls never billed.

Use cases

Batch text processing

Classification, extraction, summarising and rewriting at volume, at a low unit price that suits high-concurrency batch runs.

Image recognition and OCR

Reads text and content from screenshots, receipts and product photos; public image links and base64 both work.

Everyday chat and support

Knowledge Q&A and writing assistants that need fast, low-cost replies.

Agents and tool calling

OpenAI-format tools / tool_calls and JSON output for the execution steps of a lightweight agent.

How to choose

Choose Qwen3.8 Flash when tasks are simple, high-volume and sensitive to cost and speed; switch to Qwen3.8 Max for stronger reasoning, coding and agent work, or to the 0902 snapshot for a pinned version. All three belong to the Qwen3.8 family, so switching is just the model field.

FAQ

How do I turn on prompt caching?
cache_control is accepted: add "cache_control": {"type": "ephemeral"} to a text block in messages (write the content as [{"type": "text", "text": "...", "cache_control": {"type": "ephemeral"}}]). The upstream may serve a repeated long prefix from cache, but whether it hits is decided upstream and hits are not guaranteed. Billing follows the usage returned: the part reported in usage.prompt_tokens_details.cached_tokens settles at the cache rate, and the rest at the input rate.
Can I turn thinking off?
No. Qwen3.8 Flash always thinks before it answers, and enable_thinking, reasoning_effort or thinking_budget do not turn it off. Reasoning tokens are billed at the output rate inside usage.completion_tokens. For shorter replies, ask for concise answers in the prompt.
How do I send an image?
Put an image_url in the message content; public image links and base64 data URIs both work. In our tests it read the text in an image accurately and recognised the objects in it.
How is it different from Qwen3.8 Max?
Flash is lighter, faster and cheaper, for high-volume and simpler tasks; Max is the 3.8 flagship with stronger complex reasoning, coding and agent work. Both integrate the same way: only the model field changes.
Am I charged for failed calls?
No. Only calls that return normally are settled, at actual usage; upstream errors and timeouts never touch your balance.

Related Models

Explore other models you can integrate.

View all →