NezhaGateNezhaGate
Chat

Qwen3.8 Max 0902

⧉
qwen3.8-max-0902
Model Type Max0902Flash

Qwen3.8 Max 0902 is the 0902 snapshot of Qwen3.8 Max: the version is pinned and does not move with upstream updates, which suits workloads whose prompts are already tuned and need stable, reproducible output. Same ability and price as Qwen3.8 Max. It always thinks before it answers (the reasoning comes back in reasoning_content) and supports tool calling, streaming and image input (as base64 data URIs); a repeated long prefix hits the cache automatically and that part settles at the cache rate. Fully OpenAI-compatible: set model to qwen3.8-max-0902.

ChatPinned versionReasoningPrompt cachingOpenAI compatible

Live Test · Playground

Try out Qwen3.8 Max 0902 right here (available after login).

Input

Advanced
Web search
Memory · multi-turn

Conversation

Start a conversationType a message below to begin

About Qwen3.8 Max 0902

Qwen3.8 Max 0902 is the 0902 snapshot of Qwen3.8 Max, served on NezhaGate through the OpenAI-compatible API. The version is pinned and does not move with upstream updates, which suits workloads whose prompts are already tuned and need stable, reproducible output; ability and price match Qwen3.8 Max. It always thinks before it answers, with the reasoning in message.reasoning_content. Tool calling, streaming and image input are supported, and a repeated long prefix hits the cache automatically. Set model to qwen3.8-max-0902; pay-as-you-go, failed calls never billed.

Use cases

A pinned version in production

For live features, evaluation baselines and regression tests that need the model to behave the same way, without drift from upstream updates.

Repeated questions over long documents

Reusing the same long document or long system prompt hits the cache automatically; a prefix of about 10K tokens was 97% cached on the second call in our tests, and the cached part settles at the cache rate.

Reasoning and code

The same reasoning, coding and agent ability as Qwen3.8 Max, with OpenAI-format tool calling and JSON output.

Image understanding

Reads text and content from screenshots, receipts and charts; send images as base64 data URIs.

How to choose

Choose the 0902 snapshot for a pinned version and reproducible output, or when you reuse the same long prefix; choose Qwen3.8 Max to always get the current version; choose Qwen3.8 Flash for high volume when cost matters most. All three belong to the Qwen3.8 family, so switching is just the model field.

FAQ

How is it different from Qwen3.8 Max?
Qwen3.8 Max is the current version; 0902 is a pinned snapshot whose behaviour does not change with upstream updates. Both cost the same and integrate the same way. In our tests 0902 caches a repeated long prefix automatically, while Qwen3.8 Max currently does not.
How is caching billed?
A repeated long prefix hits the cache automatically; the cached part appears in usage.prompt_tokens_details.cached_tokens and settles at the cache rate, and the rest bills at the input rate. No extra parameter is needed.
Can I turn thinking off?
No. 0902 always thinks before it answers, and enable_thinking, reasoning_effort or thinking_budget do not change that. Reasoning tokens are billed at the output rate inside usage.completion_tokens.
How do I send an image?
Put an image_url with a base64 data URI (data:image/png;base64,...) in the message content. Convert public image links to base64 first, or use Qwen3.8 Flash, which accepts links directly.
Am I charged for failed calls?
No. Only calls that return normally are settled, at actual usage; upstream errors and timeouts never touch your balance.

Related Models

Explore other models you can integrate.

View all →