Chat, images and video,through one API
NezhaGate is an OpenAI-compatible AI model gateway: one API for GPT, Claude, Gemini, Grok, DeepSeek, Kimi, GLM and Qwen chat, plus image and video generation. Built-in API key management, balance and usage, a model market and developer docs — billed by usage, with no charge on failed requests.
Model API guides: price, measured latency, setup
One page per popular model: the live price, latency from our own weekly test runs, a cost example against the vendor's list price, and setup for Cursor, Cline, Claude Code, n8n and Open WebUI.
GPT-6 Astra API
OpenAI's GPT-6 flagship, 1.05M-token context
$1.00 per 1M input tokens and $5.00 per 1M output tokens
First token p50 1.87 s
Claude Sonnet 4.6 API
Anthropic Claude, OpenAI-compatible + Claude Code
$1.50 per 1M input tokens and $7.50 per 1M output tokens
First token p50 3.01 s
Gemini 3.1 Pro API
Google Gemini, 1M-token input, OpenAI-compatible
$0.50 per 1M input tokens and $3.00 per 1M output tokens
First token p50 9.18 s
GPT Image 2 API
OpenAI text-to-image and image editing
$0.005 per image at 1K ($0.01 at 2K, $0.02 at 4K)
Median 29.2 s per image
Nano Banana Pro API
Google Gemini 3 Pro Image, text-to-image and edits
$0.04 per image at 1K ($0.06 at 2K, $0.10 at 4K)
Median 22.9 s per image
Model market
Top models hand-picked across every category — start integrating in minutes.
GPT-5.6 Terra
Balanced everyday agentic coding model
GPT-5.5
Flagship chat and reasoning model
GPT-6 Astra
Next-generation OpenAI flagship - 1.05M context
GPT Image 2
High-quality text-to-image / image-to-image model
GPT Image 2.5 Flare
Next-gen image model - clean and smooth
Nano Banana 2
High-quality text-to-image / image-to-image model
Nano Banana Pro
Flagship text-to-image / image-to-image model
Claude Sonnet 4.6
Balanced, efficient chat and coding model
Claude Opus 5
Next-generation flagship from Anthropic
Claude Fable 5
Anthropic Fable series · narrative and long-form writing
Gemini 3.1 Pro
Next-generation flagship reasoning model from Google
Gemini 3.8 Flash
Latest Flash · adaptive thinking
Gemini 3.7 Flash
Previous Flash · adaptive thinking
Gemini 3.6 Flash
Next-generation fast thinking model · four thinking budgets
Gemini 3 Flash
Fast, low-latency, cost-efficient chat model
Gemini 2.5 Flash
Fast, low-latency, cost-efficient chat
DeepSeek V4.1 Flash
DeepSeek's current Flash model - thinking on or off
GLM-5.3
The new Z.ai flagship - coding and agents
Kimi K3
Moonshot's new flagship - long context and agents
Qwen3.7 Max
Alibaba's Qwen flagship - reasoning and coding
Grok 4.7
xAI's flagship reasoning model - long context
Veo 3.1 Not yet open
Text / image to video · multiple tiers · multiple resolutions
Seedance 2.5
Text / image to video · up to 30 seconds
Seedance 2.5 · 30s
Billed per clip · fixed 30 seconds
Wan 3.0
One endpoint, five modes · up to 30 seconds
MiniMax H3
1080P · native audio
Grok Imagine Video 1.5
Billed per clip · up to 15 seconds
Why choose NezhaGate
Built for developers and AI applications — a real platform from day one.
OpenAI-compatible, zero rework
Just swap the Base URL and Key, and your existing SDK works as-is. Copy the Base URL, Bearer Token, and Python / Node.js examples with one click — or hand the whole page to an AI.
Full control over keys, quotas, and usage
A dedicated Key and limit for each project, with real-time balance and call logs. Admins can adjust pricing, top up, and enable or disable access.
Pay-as-you-go, no charge on failure
Clear, transparent pricing settled by actual usage; upstream failures fall back automatically and are never billed to you.
NezhaGate FAQ
Common questions on access, billing, models and safety.




