🎨 GPT image generation is now cheaper: 1 credit per image, and the 4K tier renders true native 4K (3840×2160) →
NezhaGateNezhaGate
OpenAI-Compatible AI Gateway

Chat, images and video,through one API

NezhaGate is an OpenAI-compatible AI model gateway: one API for GPT, Claude, Gemini, Grok, DeepSeek, Kimi, GLM and Qwen chat, plus image and video generation. Built-in API key management, balance and usage, a model market and developer docs — billed by usage, with no charge on failed requests.

OpenAI SDK compatibleStreaming outputPay-as-you-go billingNo charge on failure
96.0%7-day success rate
OpenAIAPI format compatible
Meteredbilling · no hidden fees
26open models · always growing

Model API guides: price, measured latency, setup

One page per popular model: the live price, latency from our own weekly test runs, a cost example against the vendor's list price, and setup for Cursor, Cline, Claude Code, n8n and Open WebUI.

Model market

Top models hand-picked across every category — start integrating in minutes.

View all →
GPT-5.6 Terra Chat

GPT-5.6 Terra

gpt-5.6-terra

Balanced everyday agentic coding model

$0.20/1M in · $1.20/1M out ↓90% View →
GPT-5.5 Chat

GPT-5.5

gpt-5.5

Flagship chat and reasoning model

$0.50/1M in · $3.00/1M out ↓90% View →
GPT-6 Astra Chat

GPT-6 Astra

gpt-6-astra

Next-generation OpenAI flagship - 1.05M context

$1.00/1M in · $5.00/1M out ↓90% View →
GPT Image 2 Image

GPT Image 2

gpt-image-2

High-quality text-to-image / image-to-image model

$0.005/img and up View →
GPT Image 2.5 Flare Image

GPT Image 2.5 Flare

gpt-image-2.5-flare

Next-gen image model - clean and smooth

$0.005/img and up View →
Nano Banana 2 Image

Nano Banana 2

nano-banana-2

High-quality text-to-image / image-to-image model

$0.025/img and up View →
Nano Banana Pro Image

Nano Banana Pro

nano-banana-pro

Flagship text-to-image / image-to-image model

$0.040/img and up View →
Claude Sonnet 4.6 Chat

Claude Sonnet 4.6

claude-sonnet-4-6

Balanced, efficient chat and coding model

$1.50/1M in · $7.50/1M out ↓50% View →
Claude Opus 5 Chat

Claude Opus 5

claude-opus-5

Next-generation flagship from Anthropic

$4.00/1M in · $20.00/1M out View →
Claude Fable 5 Chat

Claude Fable 5

claude-fable-5

Anthropic Fable series · narrative and long-form writing

$8.00/1M in · $40.00/1M out View →
Gemini 3.1 Pro Chat

Gemini 3.1 Pro

gemini-3.1-pro

Next-generation flagship reasoning model from Google

$0.50/1M in · $3.00/1M out ↓75% View →
Gemini 3.8 Flash Chat

Gemini 3.8 Flash

gemini-3.8-flash

Latest Flash · adaptive thinking

$0.60/1M in · $3.60/1M out ↓7% View →
Gemini 3.7 Flash Chat

Gemini 3.7 Flash

gemini-3.7-flash

Previous Flash · adaptive thinking

$0.60/1M in · $3.60/1M out ↓7% View →
Gemini 3.6 Flash Chat

Gemini 3.6 Flash

gemini-3.6-flash

Next-generation fast thinking model · four thinking budgets

$0.60/1M in · $3.60/1M out View →
Gemini 3 Flash Chat

Gemini 3 Flash

gemini-3-flash-preview

Fast, low-latency, cost-efficient chat model

$0.30/1M in · $1.20/1M out ↓57% View →
Gemini 2.5 Flash Chat

Gemini 2.5 Flash

gemini-2.5-flash

Fast, low-latency, cost-efficient chat

$0.30/1M in · $1.20/1M out ↓46% View →
DeepSeek V4.1 Flash Chat

DeepSeek V4.1 Flash

deepseek-v4.1-flash

DeepSeek's current Flash model - thinking on or off

$0.09/1M in · $0.36/1M out ↓40% View →
GLM-5.3 Chat

GLM-5.3

glm-5.3

The new Z.ai flagship - coding and agents

$0.34/1M in · $1.25/1M out ↓73% View →
Kimi K3 Chat

Kimi K3

kimi-k3

Moonshot's new flagship - long context and agents

$1.80/1M in · $9.00/1M out ↓40% View →
Qwen3.7 Max Chat

Qwen3.7 Max

qwen3.7-max

Alibaba's Qwen flagship - reasoning and coding

$1.65/1M in · $4.85/1M out ↓35% View →
Grok 4.7 Chat

Grok 4.7

grok-4.7

xAI's flagship reasoning model - long context

$0.30/1M in · $0.90/1M out ↓85% View →
Veo 3.1 Video

Veo 3.1 Not yet open

veo-3.1

Text / image to video · multiple tiers · multiple resolutions

$0.075/clip+ View →
Seedance 2.5 Video

Seedance 2.5

seedance-2.5

Text / image to video · up to 30 seconds

$0.158/s View →
Seedance 2.5 · 30s Video

Seedance 2.5 · 30s

seedance-2.5-30s

Billed per clip · fixed 30 seconds

$1.500/clip+ View →
Wan 3.0 Video

Wan 3.0

wan3.0-video

One endpoint, five modes · up to 30 seconds

$0.120/s View →
MiniMax H3 Video

MiniMax H3

minimax-h3

1080P · native audio

$0.036/s View →
Grok Imagine Video 1.5 Video

Grok Imagine Video 1.5

grok-imagine-video-1.5

Billed per clip · up to 15 seconds

$0.300/clip+ View →

Why choose NezhaGate

Built for developers and AI applications — a real platform from day one.

Developer-first

OpenAI-compatible, zero rework

Just swap the Base URL and Key, and your existing SDK works as-is. Copy the Base URL, Bearer Token, and Python / Node.js examples with one click — or hand the whole page to an AI.

Operable

Full control over keys, quotas, and usage

A dedicated Key and limit for each project, with real-time balance and call logs. Admins can adjust pricing, top up, and enable or disable access.

Transparent

Pay-as-you-go, no charge on failure

Clear, transparent pricing settled by actual usage; upstream failures fall back automatically and are never billed to you.

FAQ

NezhaGate FAQ

Common questions on access, billing, models and safety.

What is NezhaGate?+
NezhaGate is an OpenAI-compatible AI model gateway that connects you to GPT, Claude, Gemini plus image- and video-generation models through one API and one key. Instead of integrating each provider separately, you point your Base URL at NezhaGate and call multiple flagship models with the OpenAI SDK you already use. The platform includes API key management, balance and usage tracking, a model market and developer docs, billed by usage with no charge on failed requests.
Which models are supported?+
NezhaGate currently provides OpenAI's GPT-6 Astra, the GPT-5.6 family (Sol / Terra / Luna), GPT-5.5 and the image models GPT Image 2 and GPT Image 2.5 (Flare / Sunburst); Anthropic's Claude Opus 5, Claude Fable 5 and Claude Sonnet 4.6; Google's Gemini 3.1 Pro and Gemini 3.8 / 3.7 / 3.6 / 3.5 / 3 / 2.5 Flash plus the Nano Banana 2 and Nano Banana Pro image models; DeepSeek V4.1 Flash and V4 Flash 0731; Z.ai's GLM-5.3 and GLM-5.3 Flash; Moonshot's Kimi K3; Alibaba's Qwen3.7 Max; xAI's Grok 4.7; and for video, Seedance 2.5 and the 2.0 family, Wan 3.0, MiniMax H3 and Grok Imagine Video 1.5, with new models added over time. The full, live model list and capability notes are shown on the Model Market page. Every model is called through the same endpoint and key.
How does NezhaGate bill?+
NezhaGate uses prepaid credits and pay-as-you-go billing: chat models are billed by input / output tokens and image models by count and resolution, with unit prices shown on the Pricing page and Model Market. There are no monthly plans — you pay for what you use, and the balance is drawn down by actual usage. Every call's usage and cost is visible in the console.
Am I charged for failed requests?+
Failed requests are not charged. Only calls that return successfully are billed by actual usage; upstream errors or timeouts do not consume your balance. If you see a failed record in the console, no cost is deducted for it.
How do I get started quickly with NezhaGate?+
Getting started takes three steps: register an account, create an API key in the console, then point your Base URL at NezhaGate and add the key. Because it is fully OpenAI-compatible, existing OpenAI SDK code needs almost no changes — just swap base_url and api_key. See the developer docs for details.
Are my account and API keys safe?+
An API key is shown in full exactly once, at creation; afterwards the console shows only its prefix, and you can disable or delete it at any time. Each key carries its own spend cap, requests-per-minute limit, model allowlist and IP allowlist, so a leaked key never reaches your other projects. Every call and every balance change is itemised, so your bill always reconciles with your usage. Request content is used to fulfil the call and bill it -- never to train models.
NezhaGate FAQ →