DeepSeek V4 Flash 0731
⧉DeepSeek V4 Flash 0731 is the 0731 snapshot of V4 Flash: the version is pinned and does not move with upstream updates, which suits workloads whose prompts are tuned for V4 Flash and need stable, reproducible output. Thinking is on by default (the reasoning comes back in reasoning_content) and thinking={"type":"disabled"} switches it off; tool calling and streaming are supported. Fully OpenAI-compatible: set model to deepseek-v4-flash-0731. Same price as V4.1.
Live Test · Playground
Try out DeepSeek V4 Flash 0731 right here (available after login).
Input
Advanced
Conversation
About DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is the 0731 pinned snapshot of DeepSeek V4 Flash, served on NezhaGate through the OpenAI-compatible API. The version is locked and does not change with upstream model updates, which suits workloads whose prompts are tuned for V4 Flash and that need stable, reproducible output. It thinks before it answers by default, with the reasoning in message.reasoning_content; send thinking={"type":"disabled"} to turn thinking off. Tool calling and streaming are supported. Set model to deepseek-v4-flash-0731; pay-as-you-go, failed calls never billed.
Use cases
Live services whose prompts are tuned for V4 Flash and must not change behaviour when the model is updated.
A pinned version keeps A/B evaluations and regression comparisons honest: any difference comes from your own changes.
Batch jobs that need consistent output formats can run with thinking off for direct structured results.
Keep call behaviour steady inside agent and function-calling flows, with fewer surprises from upstream upgrades.
How to choose
Choose V4.1 for DeepSeek's latest Flash model, and 0731 when you need a pinned version with reproducible output. Both cost the same, and switching is just the model field.
FAQ
What does 0731 mean?
Is it the same price as V4.1 Flash?
How do I turn thinking off?
What happens with a very small max_tokens?
Am I charged for failed calls?
Related Models
Explore other models you can integrate.
DeepSeek V4.1 Flash
DeepSeek's current Flash model - thinking on or off
GLM-5.3 Flash
The light GLM-5.3 - high volume, low cost
GLM-5.3
The new Z.ai flagship - coding and agents
Gemini 3.6 Flash
Next-generation fast thinking model · four thinking budgets