DeepSeek V4.1 Flash
⧉DeepSeek V4.1 Flash is DeepSeek's current Flash model (deepseek-flash on the official API). Thinking is on by default and comes back in reasoning_content; send thinking={"type":"disabled"} to switch it off and get a direct answer. Tool calling and streaming are supported. Fully OpenAI-compatible: set model to deepseek-v4.1-flash. One flat price around the clock with no peak-hour rate; pay-as-you-go, failed calls never billed.
Live Test · Playground
Try out DeepSeek V4.1 Flash right here (available after login).
Input
Advanced
Conversation
About DeepSeek V4.1 Flash
DeepSeek V4.1 Flash is DeepSeek's current Flash model (deepseek-flash on the official API), served on NezhaGate through the OpenAI-compatible API. It thinks before it answers by default: the reasoning comes back in message.reasoning_content and the answer always stays in content. Send thinking={"type":"disabled"} to switch thinking off for faster, cheaper answers on simple tasks. Tool calling (function calling) and streaming are supported. Set model to deepseek-v4.1-flash; one flat price around the clock with no peak-hour rate, pay-as-you-go, and failed calls are never billed.
Use cases
Support desks, knowledge Q&A and writing assistants: a low unit price and quick replies make it a fit for high-volume traffic.
With thinking on, the model works the problem through first, which suits code generation, maths and multi-step logic.
OpenAI-format tools / tool_calls plug straight into agent frameworks for retrieval, function calls and multi-step tasks.
Classification, extraction, summarising and rewriting can run with thinking off for direct, cheaper output.
How to choose
Pick V4.1 for DeepSeek's latest Flash model; pick the 0731 snapshot when your prompts are tuned for V4 Flash and you want the version pinned. For harder coding and agent work, compare GLM-5.3 - switching is just the model field.
FAQ
How is it different from DeepSeek V4 Flash 0731?
How do I turn thinking off?
What happens with a very small max_tokens?
Does it support tool calling and streaming?
Am I charged for failed calls?
Related Models
Explore other models you can integrate.
DeepSeek V4 Flash 0731
Pinned V4 Flash snapshot - a version that stays put
GLM-5.3
The new Z.ai flagship - coding and agents
GLM-5.3 Flash
The light GLM-5.3 - high volume, low cost
Gemini 3.6 Flash
Next-generation fast thinking model · four thinking budgets
GPT-6 Luna
The most efficient GPT-6, built for high-volume work