NezhaGate
Chat

Gemini 3.6 Flash

gemini-3.6-flash
Model Type MediumHighLowTiered

Gemini 3.6 Flash is the next-generation fast thinking model from Google: 1M-token context, up to 64K output, image input. The standard (Medium) tier has a balanced thinking budget for everyday chat and agent workflows. Fully OpenAI-API compatible: set model to gemini-3.6-flash.

ChatThinking1M contextOpenAI compatible

Live Test · Playground

Try out Gemini 3.6 Flash right here (available after login)。

Input

Advanced
Web search
Memory · multi-turn

Conversation

Start a conversationType a message below to begin

About Gemini 3.6 Flash

Gemini 3.6 Flash is Google's new-generation fast thinking model: it adds an explicit thinking (reasoning) step at Flash-tier latency and cost, handles a 1,000,000-token context with up to 64K tokens of output, and accepts image input. NezhaGate exposes four thinking-budget tiers: the standard gemini-3.6-flash, the deeper gemini-3.6-flash-high, the fastest gemini-3.6-flash-low, and gemini-3.6-flash-tiered, where the model decides per request how much to think. All four cost the same per token — how much a tier thinks simply shows up as output tokens. It is served through an OpenAI-compatible endpoint: point base_url at NezhaGate and set model to "gemini-3.6-flash" to reuse your existing OpenAI SDK code, with pay-as-you-go billing where failed requests are not charged.

Use cases

Everyday chat and agents that need to think

The standard tier balances speed and reasoning, fitting support bots, assistants and agent workflows that must answer quickly yet still reason.

Million-token document work

The 1M context window swallows a large repository, a long report or a stack of contracts in one pass for summarising, retrieval and cross-checking.

Mixed text and image understanding

Image input supports screenshot Q&A, reading forms and charts, and other multimodal analysis.

Workloads of uneven difficulty

The tiered variant lets the model choose its own thinking budget per request — quick answers for easy questions, more deliberation for hard ones — which suits production traffic with a wide difficulty spread.

How to choose

Use the standard gemini-3.6-flash by default; switch to gemini-3.6-flash-high for multi-step reasoning; use gemini-3.6-flash-low for bulk classification, extraction and rewriting, where it is both fastest and cheapest in practice; and pick gemini-3.6-flash-tiered when request difficulty is unpredictable and you would rather let the model decide. All four share one price, so changing tier is just changing the model field. Step up to Gemini 3.1 Pro when you need more reasoning depth.

FAQ

What is the difference between the four Gemini 3.6 Flash tiers?
Only the thinking budget: low thinks least and answers fastest, the standard (unsuffixed) id sits in the middle, high thinks most and suits multi-step reasoning, and tiered lets the model decide per request. They share the same underlying model, context window and multimodal support.
Why do all four tiers cost the same?
Because the difference already shows up in usage: thinking tokens are billed as output tokens, so the high tier naturally consumes more output tokens than the low tier. Adding a separate per-token premium on top would charge for the same thing twice, so all four share one rate and your actual bill follows how much the tier really thought.
How large is the context window?
About 1,000,000 input tokens with up to 64K output tokens per call, identical across all four tiers.
Does it support image input and streaming?
Both. Pass images using OpenAI's image_url format, and set stream=true in the request to receive the answer token by token.
How do I call it with the OpenAI SDK?
Point base_url at NezhaGate, supply your NezhaGate API key, and set model to "gemini-3.6-flash" (or one of the suffixed tier ids). Everything else matches your existing OpenAI code.

Related Models

Explore other models you can integrate.

View all →
GPT-5.6 Sol Chat

GPT-5.6 Sol

gpt-5.6-sol

Frontier flagship for agentic coding

$2.00/1M in · $12.00/1M out ↓75% View →
GPT-5.6 Terra Chat

GPT-5.6 Terra

gpt-5.6-terra

Balanced everyday agentic coding model

$1.20/1M in · $7.00/1M out ↓77% View →
GPT-5.6 Luna Chat

GPT-5.6 Luna

gpt-5.6-luna

Fast, economical agentic coding model

$0.80/1M in · $4.80/1M out ↓68% View →
GPT-5.5 Chat

GPT-5.5

gpt-5.5

Flagship chat and reasoning model

$0.70/1M in · $4.20/1M out ↓86% View →
GPT-6 Astra Chat

GPT-6 Astra

gpt-6-astra

Next-generation OpenAI flagship - 1.05M context

$2.80/1M in · $14.00/1M out ↓72% View →
GPT Image 2 Image

GPT Image 2

gpt-image-2

High-quality text-to-image / image-to-image model

$0.015/img and up View →
GPT Image 2.5 Flare Image

GPT Image 2.5 Flare

gpt-image-2.5-flare

Next-gen image model - clean and smooth

$0.015/img and up View →
GPT Image 2.5 Sunburst Image

GPT Image 2.5 Sunburst

gpt-image-2.5-sunburst

Next-gen image model - richer texture

$0.015/img and up View →
Nano Banana 2 Image

Nano Banana 2

nano-banana-2

High-quality text-to-image / image-to-image model

$0.025/img and up View →
Nano Banana Pro Image

Nano Banana Pro

nano-banana-pro

Flagship text-to-image / image-to-image model

$0.040/img and up View →
Claude Sonnet 4.6 Chat

Claude Sonnet 4.6

claude-sonnet-4-6

Balanced, efficient chat and coding model

$1.50/1M in · $7.50/1M out ↓50% View →
Claude Opus 5 Chat

Claude Opus 5

claude-opus-5

Next-generation flagship from Anthropic

$4.00/1M in · $20.00/1M out View →
Claude Fable 5 Chat

Claude Fable 5

claude-fable-5

Anthropic Fable series · narrative and long-form writing

$8.00/1M in · $40.00/1M out View →
Gemini 3.1 Pro Chat

Gemini 3.1 Pro

gemini-3.1-pro

Next-generation flagship reasoning model from Google

$0.50/1M in · $3.00/1M out ↓75% View →
Gemini 3.8 Flash Chat

Gemini 3.8 Flash

gemini-3.8-flash

Latest Flash · adaptive thinking

$0.60/1M in · $3.60/1M out ↓7% View →
Gemini 3.7 Flash Chat

Gemini 3.7 Flash

gemini-3.7-flash

Previous Flash · adaptive thinking

$0.60/1M in · $3.60/1M out ↓7% View →
Gemini 3.6 Flash High Chat

Gemini 3.6 Flash High

gemini-3.6-flash-high

Deep thinking tier

$0.60/1M in · $3.60/1M out View →
Gemini 3.6 Flash Low Chat

Gemini 3.6 Flash Low

gemini-3.6-flash-low

Fastest, lowest-cost tier

$0.60/1M in · $3.60/1M out View →
Gemini 3.6 Flash Tiered Chat

Gemini 3.6 Flash Tiered

gemini-3.6-flash-tiered

Auto-tiered thinking

$0.60/1M in · $3.60/1M out View →
Gemini 3 Flash Chat

Gemini 3 Flash

gemini-3-flash-preview

Fast, low-latency, cost-efficient chat model

$0.30/1M in · $1.20/1M out ↓57% View →
Gemini 3.5 Flash Chat

Gemini 3.5 Flash

gemini-3.5-flash

Next-generation fast thinking model

$0.45/1M in · $2.70/1M out ↓70% View →
Gemini 2.5 Flash Chat

Gemini 2.5 Flash

gemini-2.5-flash

Fast, low-latency, cost-efficient chat

$0.30/1M in · $1.20/1M out ↓46% View →
Veo 3.1 Video

Veo 3.1

veo-3.1

Text / image to video · multiple tiers · multiple resolutions

$0.075/clip+ View →
Seedance 2.5 Video

Seedance 2.5

seedance-2.5

Text / image to video · up to 30 seconds

$0.158/s View →
Seedance 2.0 Video

Seedance 2.0

seedance-2.0

Text / image to video · 5-15 seconds

$0.150/s View →
Seedance 2.0 Fast Video

Seedance 2.0 Fast

seedance-2.0-fast

Low latency · lowest cost

$0.100/s View →
Seedance 2.0 Mini Video

Seedance 2.0 Mini

seedance-2.0-mini

Entry tier · lowest cost

$0.066/s View →
Wan 3.0 Video

Wan 3.0

wan3.0-video

One endpoint, five modes · up to 30 seconds

$0.120/s View →
Wan 3.0 Prime Video

Wan 3.0 Prime

wan3.0-video-prime

Same capabilities · several times faster

$0.160/s View →
MiniMax H3 Video

MiniMax H3

minimax-h3

1080P · native audio

$0.036/s View →
Grok Imagine Video 1.5 Video

Grok Imagine Video 1.5

grok-imagine-video-1.5

Billed per clip · up to 15 seconds

$0.300/clip+ View →