Qwen3.8 Flash
⧉Qwen3.8 Flash is the light, fast member of the Qwen3.8 family, with a low unit price and high throughput for batch jobs and high concurrency; short prompts came back in 2.5-6 seconds in our tests. It always thinks before it answers (the reasoning comes back in reasoning_content) and supports tool calling, JSON output, streaming and image input (public image links or base64). cache_control on a text block is accepted, and the upstream may serve a repeated long prefix from cache; whether it hits is decided upstream, and billing follows the returned usage (a reported cache hit settles at the cache rate). Fully OpenAI-compatible: set model to qwen3.8-flash. Pay-as-you-go, failed calls never billed.
Live Test · Playground
Try out Qwen3.8 Flash right here (available after login).
Input
Advanced
Conversation
About Qwen3.8 Flash
Qwen3.8 Flash is the light, fast member of Alibaba Cloud's Qwen3.8 family, served on NezhaGate through the OpenAI-compatible API. It has a low unit price and high throughput for batch jobs and high concurrency; short prompts came back in 2.5-6 seconds in our tests. It always thinks before it answers, with the reasoning in message.reasoning_content. Tool calling, JSON output, streaming and image input (public links or base64) are supported; cache_control is accepted too, and the upstream may serve a repeated long prefix from cache: whether it hits is decided upstream, and billing follows the returned usage (a reported cache hit settles at the cache rate). Set model to qwen3.8-flash; pay-as-you-go, failed calls never billed.
Use cases
Classification, extraction, summarising and rewriting at volume, at a low unit price that suits high-concurrency batch runs.
Reads text and content from screenshots, receipts and product photos; public image links and base64 both work.
Knowledge Q&A and writing assistants that need fast, low-cost replies.
OpenAI-format tools / tool_calls and JSON output for the execution steps of a lightweight agent.
How to choose
Choose Qwen3.8 Flash when tasks are simple, high-volume and sensitive to cost and speed; switch to Qwen3.8 Max for stronger reasoning, coding and agent work, or to the 0902 snapshot for a pinned version. All three belong to the Qwen3.8 family, so switching is just the model field.
FAQ
How do I turn on prompt caching?
Can I turn thinking off?
How do I send an image?
How is it different from Qwen3.8 Max?
Am I charged for failed calls?
Related Models
Explore other models you can integrate.
Qwen3.8 Max
Alibaba's Qwen 3.8 flagship - reasoning, coding and agents
Qwen3.8 Max 0902
Pinned Qwen3.8 Max snapshot - a version that stays put
DeepSeek V4.1 Flash
DeepSeek's current Flash model - thinking on or off
GLM-5.3 Flash
The light GLM-5.3 - high volume, low cost
Gemini 3.6 Flash
Next-generation fast thinking model · four thinking budgets