Gemini 3.6 Flash
⧉Gemini 3.6 Flash is the next-generation fast thinking model from Google: 1M-token context, up to 64K output, image input. The standard (Medium) tier has a balanced thinking budget for everyday chat and agent workflows. Fully OpenAI-API compatible: set model to gemini-3.6-flash.
Live Test · Playground
Try out Gemini 3.6 Flash right here (available after login)。
Input
Advanced
Conversation
About Gemini 3.6 Flash
Gemini 3.6 Flash is Google's new-generation fast thinking model: it adds an explicit thinking (reasoning) step at Flash-tier latency and cost, handles a 1,000,000-token context with up to 64K tokens of output, and accepts image input. NezhaGate exposes four thinking-budget tiers: the standard gemini-3.6-flash, the deeper gemini-3.6-flash-high, the fastest gemini-3.6-flash-low, and gemini-3.6-flash-tiered, where the model decides per request how much to think. All four cost the same per token — how much a tier thinks simply shows up as output tokens. It is served through an OpenAI-compatible endpoint: point base_url at NezhaGate and set model to "gemini-3.6-flash" to reuse your existing OpenAI SDK code, with pay-as-you-go billing where failed requests are not charged.
Use cases
The standard tier balances speed and reasoning, fitting support bots, assistants and agent workflows that must answer quickly yet still reason.
The 1M context window swallows a large repository, a long report or a stack of contracts in one pass for summarising, retrieval and cross-checking.
Image input supports screenshot Q&A, reading forms and charts, and other multimodal analysis.
The tiered variant lets the model choose its own thinking budget per request — quick answers for easy questions, more deliberation for hard ones — which suits production traffic with a wide difficulty spread.
How to choose
Use the standard gemini-3.6-flash by default; switch to gemini-3.6-flash-high for multi-step reasoning; use gemini-3.6-flash-low for bulk classification, extraction and rewriting, where it is both fastest and cheapest in practice; and pick gemini-3.6-flash-tiered when request difficulty is unpredictable and you would rather let the model decide. All four share one price, so changing tier is just changing the model field. Step up to Gemini 3.1 Pro when you need more reasoning depth.
FAQ
What is the difference between the four Gemini 3.6 Flash tiers?
Why do all four tiers cost the same?
How large is the context window?
Does it support image input and streaming?
How do I call it with the OpenAI SDK?
Related Models
Explore other models you can integrate.
GPT-5.6 Sol
Frontier flagship for agentic coding
GPT-5.6 Terra
Balanced everyday agentic coding model
GPT-5.6 Luna
Fast, economical agentic coding model
GPT-5.5
Flagship chat and reasoning model
GPT-6 Astra
Next-generation OpenAI flagship - 1.05M context
GPT Image 2
High-quality text-to-image / image-to-image model
GPT Image 2.5 Flare
Next-gen image model - clean and smooth
GPT Image 2.5 Sunburst
Next-gen image model - richer texture
Nano Banana 2
High-quality text-to-image / image-to-image model
Nano Banana Pro
Flagship text-to-image / image-to-image model
Claude Sonnet 4.6
Balanced, efficient chat and coding model
Claude Opus 5
Next-generation flagship from Anthropic
Claude Fable 5
Anthropic Fable series · narrative and long-form writing
Gemini 3.1 Pro
Next-generation flagship reasoning model from Google
Gemini 3.8 Flash
Latest Flash · adaptive thinking
Gemini 3.7 Flash
Previous Flash · adaptive thinking
Gemini 3.6 Flash High
Deep thinking tier
Gemini 3.6 Flash Low
Fastest, lowest-cost tier
Gemini 3.6 Flash Tiered
Auto-tiered thinking
Gemini 3 Flash
Fast, low-latency, cost-efficient chat model
Gemini 3.5 Flash
Next-generation fast thinking model
Gemini 2.5 Flash
Fast, low-latency, cost-efficient chat
Veo 3.1
Text / image to video · multiple tiers · multiple resolutions
Seedance 2.5
Text / image to video · up to 30 seconds
Seedance 2.0
Text / image to video · 5-15 seconds
Seedance 2.0 Fast
Low latency · lowest cost
Seedance 2.0 Mini
Entry tier · lowest cost
Wan 3.0
One endpoint, five modes · up to 30 seconds
Wan 3.0 Prime
Same capabilities · several times faster
MiniMax H3
1080P · native audio
Grok Imagine Video 1.5
Billed per clip · up to 15 seconds