NezhaGate
Chat

Gemini 3.6 Flash

gemini-3.6-flash
Model Type MediumHighLowTiered

Gemini 3.6 Flash 是 Google 新一代高速思考模型,1M token 上下文、最高 64K 输出,支持图片输入。标准档(Medium)思考预算均衡,适合日常对话与 Agent 工作流。完全兼容 OpenAI 接口,把 model 改成 gemini-3.6-flash 即可接入。

文本对话思考1M 上下文OpenAI 兼容

Live Test · Playground

Try out Gemini 3.6 Flash right here (available after login)。

Input

Advanced
联网搜索 · Web search
Memory · multi-turn

Conversation

Start a conversationType a message below to begin

About Gemini 3.6 Flash

Gemini 3.6 Flash is Google's new-generation fast thinking model: it adds an explicit thinking (reasoning) step at Flash-tier latency and cost, handles a 1,000,000-token context with up to 64K tokens of output, and accepts image input. NezhaGate exposes four thinking-budget tiers: the standard gemini-3.6-flash, the deeper gemini-3.6-flash-high, the fastest gemini-3.6-flash-low, and gemini-3.6-flash-tiered, where the model decides per request how much to think. All four cost the same per token — how much a tier thinks simply shows up as output tokens. It is served through an OpenAI-compatible endpoint: point base_url at NezhaGate and set model to "gemini-3.6-flash" to reuse your existing OpenAI SDK code, with pay-as-you-go billing where failed requests are not charged.

Use cases

Everyday chat and agents that need to think

The standard tier balances speed and reasoning, fitting support bots, assistants and agent workflows that must answer quickly yet still reason.

Million-token document work

The 1M context window swallows a large repository, a long report or a stack of contracts in one pass for summarising, retrieval and cross-checking.

Mixed text and image understanding

Image input supports screenshot Q&A, reading forms and charts, and other multimodal analysis.

Workloads of uneven difficulty

The tiered variant lets the model choose its own thinking budget per request — quick answers for easy questions, more deliberation for hard ones — which suits production traffic with a wide difficulty spread.

How to choose

Use the standard gemini-3.6-flash by default; switch to gemini-3.6-flash-high for multi-step reasoning; use gemini-3.6-flash-low for bulk classification, extraction and rewriting, where it is both fastest and cheapest in practice; and pick gemini-3.6-flash-tiered when request difficulty is unpredictable and you would rather let the model decide. All four share one price, so changing tier is just changing the model field. Step up to Gemini 3.1 Pro when you need more reasoning depth.

FAQ

What is the difference between the four Gemini 3.6 Flash tiers?
Only the thinking budget: low thinks least and answers fastest, the standard (unsuffixed) id sits in the middle, high thinks most and suits multi-step reasoning, and tiered lets the model decide per request. They share the same underlying model, context window and multimodal support.
Why do all four tiers cost the same?
Because the difference already shows up in usage: thinking tokens are billed as output tokens, so the high tier naturally consumes more output tokens than the low tier. Adding a separate per-token premium on top would charge for the same thing twice, so all four share one rate and your actual bill follows how much the tier really thought.
How large is the context window?
About 1,000,000 input tokens with up to 64K output tokens per call, identical across all four tiers.
Does it support image input and streaming?
Both. Pass images using OpenAI's image_url format, and set stream=true in the request to receive the answer token by token.
How do I call it with the OpenAI SDK?
Point base_url at NezhaGate, supply your NezhaGate API key, and set model to "gemini-3.6-flash" (or one of the suffixed tier ids). Everything else matches your existing OpenAI code.

Related Models

Explore other models you can integrate.

View all →
GPT-5.6 Sol Chat

GPT-5.6 Sol

gpt-5.6-sol

前沿旗舰 Agentic 编程模型

$2.00/1M tokens ↓75% View →
GPT-5.6 Terra Chat

GPT-5.6 Terra

gpt-5.6-terra

日常均衡 Agentic 编程模型

$1.20/1M tokens ↓77% View →
GPT-5.6 Luna Chat

GPT-5.6 Luna

gpt-5.6-luna

快而省的 Agentic 编程模型

$0.80/1M tokens ↓68% View →
GPT-5.5 Chat

GPT-5.5

gpt-5.5

旗舰对话与推理模型

$0.70/1M tokens ↓86% View →
GPT Image 2 Image

GPT Image 2

gpt-image-2

高质量文生图 / 图生图模型

$0.015/img and up View →
Nano Banana 2 Image

Nano Banana 2

nano-banana-2

高质量文生图 / 图生图模型

$0.025/img and up View →
Nano Banana Pro Image

Nano Banana Pro

nano-banana-pro

旗舰级文生图 / 图生图模型

$0.040/img and up View →
Claude Sonnet 4.6 Chat

Claude Sonnet 4.6

claude-sonnet-4-6

均衡高效的对话 / 代码模型

$1.50/1M tokens ↓50% View →
Gemini 3.1 Pro Chat

Gemini 3.1 Pro

gemini-3.1-pro

Google 新一代旗舰推理模型

$0.50/1M tokens ↓75% View →
Gemini 3.6 Flash High Chat

Gemini 3.6 Flash High

gemini-3.6-flash-high

深度思考档

$0.60/1M tokens View →
Gemini 3.6 Flash Low Chat

Gemini 3.6 Flash Low

gemini-3.6-flash-low

极速低耗档

$0.60/1M tokens View →
Gemini 3.6 Flash Tiered Chat

Gemini 3.6 Flash Tiered

gemini-3.6-flash-tiered

自动调档

$0.60/1M tokens View →
Gemini 3 Flash Chat

Gemini 3 Flash

gemini-3-flash-preview

高速低延迟,高性价比对话模型

$0.30/1M tokens ↓60% View →
Gemini 3.5 Flash Chat

Gemini 3.5 Flash

gemini-3.5-flash

新一代高速思考模型

$0.45/1M tokens ↓70% View →
Gemini 2.5 Flash Chat

Gemini 2.5 Flash

gemini-2.5-flash

高速低延迟,高性价比对话

$0.30/1M tokens ↓52% View →
Veo 3.1 Video

Veo 3.1

veo-3.1

文生 / 图生视频 · 多档 · 多分辨率

$0.075/clip+ View →