Live Test · Playground
Try out Gemini 3.6 Flash right here (available after login)。
Input
Advanced
Conversation
About Gemini 3.6 Flash
Gemini 3.6 Flash is Google's new-generation fast thinking model: it adds an explicit thinking (reasoning) step at Flash-tier latency and cost, handles a 1,000,000-token context with up to 64K tokens of output, and accepts image input. NezhaGate exposes four thinking-budget tiers: the standard gemini-3.6-flash, the deeper gemini-3.6-flash-high, the fastest gemini-3.6-flash-low, and gemini-3.6-flash-tiered, where the model decides per request how much to think. All four cost the same per token — how much a tier thinks simply shows up as output tokens. It is served through an OpenAI-compatible endpoint: point base_url at NezhaGate and set model to "gemini-3.6-flash" to reuse your existing OpenAI SDK code, with pay-as-you-go billing where failed requests are not charged.
Use cases
The standard tier balances speed and reasoning, fitting support bots, assistants and agent workflows that must answer quickly yet still reason.
The 1M context window swallows a large repository, a long report or a stack of contracts in one pass for summarising, retrieval and cross-checking.
Image input supports screenshot Q&A, reading forms and charts, and other multimodal analysis.
The tiered variant lets the model choose its own thinking budget per request — quick answers for easy questions, more deliberation for hard ones — which suits production traffic with a wide difficulty spread.
How to choose
Use the standard gemini-3.6-flash by default; switch to gemini-3.6-flash-high for multi-step reasoning; use gemini-3.6-flash-low for bulk classification, extraction and rewriting, where it is both fastest and cheapest in practice; and pick gemini-3.6-flash-tiered when request difficulty is unpredictable and you would rather let the model decide. All four share one price, so changing tier is just changing the model field. Step up to Gemini 3.1 Pro when you need more reasoning depth.
FAQ
What is the difference between the four Gemini 3.6 Flash tiers?
Why do all four tiers cost the same?
How large is the context window?
Does it support image input and streaming?
How do I call it with the OpenAI SDK?
Related Models
Explore other models you can integrate.
GPT-5.6 Sol
前沿旗舰 Agentic 编程模型
GPT-5.6 Terra
日常均衡 Agentic 编程模型
GPT-5.6 Luna
快而省的 Agentic 编程模型
GPT-5.5
旗舰对话与推理模型
GPT Image 2
高质量文生图 / 图生图模型
Nano Banana 2
高质量文生图 / 图生图模型
Nano Banana Pro
旗舰级文生图 / 图生图模型
Claude Sonnet 4.6
均衡高效的对话 / 代码模型
Gemini 3.1 Pro
Google 新一代旗舰推理模型
Gemini 3.6 Flash High
深度思考档
Gemini 3.6 Flash Low
极速低耗档
Gemini 3.6 Flash Tiered
自动调档
Gemini 3 Flash
高速低延迟,高性价比对话模型
Gemini 3.5 Flash
新一代高速思考模型
Gemini 2.5 Flash
高速低延迟,高性价比对话
Veo 3.1
文生 / 图生视频 · 多档 · 多分辨率