Qwen3.8 Max 0902
⧉Qwen3.8 Max 0902 is the 0902 snapshot of Qwen3.8 Max: the version is pinned and does not move with upstream updates, which suits workloads whose prompts are already tuned and need stable, reproducible output. Same ability and price as Qwen3.8 Max. It always thinks before it answers (the reasoning comes back in reasoning_content) and supports tool calling, streaming and image input (as base64 data URIs); a repeated long prefix hits the cache automatically and that part settles at the cache rate. Fully OpenAI-compatible: set model to qwen3.8-max-0902.
Live Test · Playground
Try out Qwen3.8 Max 0902 right here (available after login).
Input
Advanced
Conversation
About Qwen3.8 Max 0902
Qwen3.8 Max 0902 is the 0902 snapshot of Qwen3.8 Max, served on NezhaGate through the OpenAI-compatible API. The version is pinned and does not move with upstream updates, which suits workloads whose prompts are already tuned and need stable, reproducible output; ability and price match Qwen3.8 Max. It always thinks before it answers, with the reasoning in message.reasoning_content. Tool calling, streaming and image input are supported, and a repeated long prefix hits the cache automatically. Set model to qwen3.8-max-0902; pay-as-you-go, failed calls never billed.
Use cases
For live features, evaluation baselines and regression tests that need the model to behave the same way, without drift from upstream updates.
Reusing the same long document or long system prompt hits the cache automatically; a prefix of about 10K tokens was 97% cached on the second call in our tests, and the cached part settles at the cache rate.
The same reasoning, coding and agent ability as Qwen3.8 Max, with OpenAI-format tool calling and JSON output.
Reads text and content from screenshots, receipts and charts; send images as base64 data URIs.
How to choose
Choose the 0902 snapshot for a pinned version and reproducible output, or when you reuse the same long prefix; choose Qwen3.8 Max to always get the current version; choose Qwen3.8 Flash for high volume when cost matters most. All three belong to the Qwen3.8 family, so switching is just the model field.
FAQ
How is it different from Qwen3.8 Max?
How is caching billed?
Can I turn thinking off?
How do I send an image?
Am I charged for failed calls?
Related Models
Explore other models you can integrate.
Qwen3.8 Max
Alibaba's Qwen 3.8 flagship - reasoning, coding and agents
Qwen3.8 Flash
The light, fast Qwen3.8 - high volume, low cost
Qwen3.7 Max
Alibaba's Qwen flagship - reasoning and coding
Kimi K3
Moonshot's new flagship - long context and agents
GLM-5.3
The new Z.ai flagship - coding and agents