# NezhaGate - complete reference (llms-full.txt) > NezhaGate (a.k.a. NezhaAPI) is an OpenAI-compatible AI model gateway: one endpoint and one API key for GPT, Claude, Gemini, Grok, DeepSeek, Kimi, GLM and Qwen chat plus image and video generation (Seedance, Veo, Wan, MiniMax H3, Nano Banana, GPT Image). Pay as you go in USD; failed requests are never billed. Base URL: https://nezhagate.com/v1. Last updated 2026-10-03. Every price below is read live from the billing table at the moment this file is generated. Official list prices carry the date we checked them; a saving is shown only where we are below the official price. Measured figures come from real customer calls and appear only when the sample is large enough. ``` base_url: https://nezhagate.com/v1 Authorization: Bearer ``` ## Key facts - API: OpenAI-compatible. POST https://nezhagate.com/v1/chat/completions and /v1/responses for chat; /v1/images/generations and /v1/images/edits for images; /v1/videos/generations for video. Images and videos are async jobs: submit, then poll GET /v1/images/jobs/{id} or /v1/videos/jobs/{id}. Claude models are also served on the Anthropic Messages API at https://nezhagate.com/anthropic/v1/messages. - Billing: chat per 1M input / output tokens, cache reads at 10% of the input price; images per picture by resolution; video per second of output (Grok Imagine Video 1.5, Seedance 2.5 · 30s, Veo 3.1 per clip). Failed requests cost nothing; image and video jobs hold credit at submit and refund it in full if they fail. - Top-up: WeChat Pay from CNY 10, USDT and bank card from $5; $1 = 200 credits; credits never expire; unused top-ups are refundable within 7 days ($0.40 fee). Larger packs add credit: $50 -> $55 (+10%), $500 -> $575 (+15%), $1,250 -> $1,500 (+20%). - Availability: the public API answered 100% of external probes over the last 7 days (one probe every 2 minutes); live per-model status at https://nezhagate.com/en/status. - Failover: the main models have several supply lines; a request that errors, times out or finds a line full moves to the next line of the same model before any output is streamed. - Rate limits: no default per-account limit; when every line of a model is full the API answers 429 with Retry-After. Per-key RPM, daily budget, total quota, model and IP allow-lists are self-service. - Data: call logs stay 60 days in the console, their text content is deleted after 30 days, generated images and videos after 60 days, uploaded references after 7 days. Content is never used for training. - Support: support@nezhagate.com, Telegram https://t.me/+-erBoH9-AYY4MTM1 and WeChat; Beijing time 10:00-19:00, replies within 1 hour. ## Chat models (live prices, USD per 1M tokens) - [GPT-5.6 Sol](https://nezhagate.com/en/model/gpt-5.6-sol): OpenAI. input $0.40 / output $2.00, cache read $0.04. official $4.00 / $20.00 (vendor list price, checked 2026-10-02), save 90%. Frontier reasoning, agentic coding, long-horizon tasks, structured output. ~272K tokens (GPT-5 family context window; system-prompt overhead counts toward it). docs: https://nezhagate.com/en/docs/gpt-5.6-sol. - [GPT-5.6 Terra](https://nezhagate.com/en/model/gpt-5.6-terra): OpenAI. input $0.20 / output $1.20, cache read $0.02. official $2.00 / $12.00 (vendor list price, checked 2026-10-02), save 90%. Everyday chat, agentic coding, reasoning, structured output. ~272K tokens (GPT-5 family context window; system-prompt overhead counts toward it). measured on real customer calls over the last 7 days: 100% success. docs: https://nezhagate.com/en/docs/gpt-5.6-terra. - [GPT-5.6 Luna](https://nezhagate.com/en/model/gpt-5.6-luna): OpenAI. input $0.07 / output $0.40, cache read $0.007. official $0.20 / $1.20 (vendor list price, checked 2026-10-02), save 65% / 66%. High-speed chat, agentic coding, high-volume low-latency, structured output. ~272K tokens (GPT-5 family context window; system-prompt overhead counts toward it). docs: https://nezhagate.com/en/docs/gpt-5.6-luna. - [GPT-5.5](https://nezhagate.com/en/model/gpt-5.5): OpenAI. input $0.50 / output $3.00, cache read $0.05. official $5.00 / $30.00 (vendor list price, checked 2026-10-02), save 90%. Chat, reasoning, agents, structured output. ~272K tokens (GPT-5 family context window; system-prompt overhead counts toward it). measured on real customer calls over the last 30 days: 100% success. docs: https://nezhagate.com/en/docs/gpt-5.5. - [GPT-6 Astra](https://nezhagate.com/en/model/gpt-6-astra): OpenAI. input $1.00 / output $5.00, cache read $0.10. official $10.00 / $50.00 (vendor list price, checked 2026-10-02), save 90%. Deep reasoning, agentic coding, very long context, image input, structured output. 1,050,000 tokens (OpenAI published figure); up to 128,000 output tokens per call. System-prompt overhead counts toward it. measured on real customer calls over the last 7 days: 96% success. first token p50 2.1 s (weekly test, 2026-09-29). docs: https://nezhagate.com/en/docs/gpt-6-astra. - [GPT-6.1 Sol](https://nezhagate.com/en/model/gpt-6.1-sol): OpenAI. input $0.40 / output $2.00, cache read $0.04. official $2.00 / $10.00 (vendor list price, checked 2026-10-02), save 80%. Agentic coding, complex coding, professional document and analysis work, reasoning, very long context, image input, structured output. 1,050,000 tokens (OpenAI published figure); up to 128,000 output tokens per call. System-prompt overhead counts toward it. docs: https://nezhagate.com/en/docs/gpt-6.1-sol. - [GPT-6 Sol](https://nezhagate.com/en/model/gpt-6-sol): OpenAI. input $0.20 / output $1.00, cache read $0.02. official $2.00 / $10.00 (vendor list price, checked 2026-10-02), save 90%. Complex coding, agentic workflows, reasoning, very long context, image input, structured output. 1,050,000 tokens (OpenAI published figure); up to 128,000 output tokens per call. System-prompt overhead counts toward it. docs: https://nezhagate.com/en/docs/gpt-6-sol. - [GPT-6 Luna](https://nezhagate.com/en/model/gpt-6-luna): OpenAI. input $0.07 / output $0.40, cache read $0.007. official $0.10 / $0.50 (vendor list price, checked 2026-10-02), save 30% / 20%. High-volume chat, low-latency agent steps, classification and extraction, very long context, structured output. 1,050,000 tokens (OpenAI published figure); up to 128,000 output tokens per call. System-prompt overhead counts toward it. docs: https://nezhagate.com/en/docs/gpt-6-luna. - [Claude Sonnet 4.6](https://nezhagate.com/en/model/claude-sonnet-4-6): Anthropic. input $1.50 / output $7.50, cache read $0.15. official $3.00 / $15.00 (vendor list price, checked 2026-10-02), save 50%. Chat, code, reasoning, long context. ~200K tokens. first token p50 2.0 s (weekly test, 2026-09-29). docs: https://nezhagate.com/en/docs/claude-sonnet-4-6. - [Claude Opus 5](https://nezhagate.com/en/model/claude-opus-5): Anthropic. input $4.00 / output $20.00, cache read $0.40. official $5.00 / $25.00 (vendor list price, checked 2026-10-02), save 20%. Deep reasoning, code, agents, long context, image understanding. ~200K tokens. measured on real customer calls over the last 30 days: 100% success. docs: https://nezhagate.com/en/docs/claude-opus-5. - [Claude Fable 5](https://nezhagate.com/en/model/claude-fable-5): Anthropic. input $8.00 / output $40.00, cache read $0.80. official $10.00 / $50.00 (vendor list price, checked 2026-10-02), save 20%. Chinese writing, narrative creation, long-form generation, chat, code, image understanding. ~200K tokens. docs: https://nezhagate.com/en/docs/claude-fable-5. - [Claude Sonnet 5](https://nezhagate.com/en/model/claude-sonnet-5): Anthropic. input $1.00 / output $5.00, cache read $0.10. official $2.00 / $10.00 (vendor list price, checked 2026-10-02), save 50%. Code, agents, reasoning, long context, image understanding. ~200K tokens. docs: https://nezhagate.com/en/docs/claude-sonnet-5. - [Gemini 3.1 Pro](https://nezhagate.com/en/model/gemini-3.1-pro): Google. input $0.50 / output $3.00, cache read $0.05. official $2.00 / $12.00 (vendor list price, checked 2026-10-02), save 75%. Chat, reasoning, very long context, multimodal. ~1,000,000 tokens (million-token long context). measured on real customer calls over the last 7 days: 100% success. first token p50 9.7 s (weekly test, 2026-09-29). docs: https://nezhagate.com/en/docs/gemini-3.1-pro. - [Gemini 3.8 Flash](https://nezhagate.com/en/model/gemini-3.8-flash): Google. input $0.60 / output $3.60, cache read $0.06. official $0.75 / $3.75 (vendor list price, checked 2026-10-02), save 20% / 4%. Chat, reasoning, adaptive thinking, image input, very long context. ~1,000,000 tokens (million-token long context). measured on real customer calls over the last 30 days: 98.6% success. docs: https://nezhagate.com/en/docs/gemini-3.8-flash. - [Gemini 3.7 Flash](https://nezhagate.com/en/model/gemini-3.7-flash): Google. input $0.60 / output $3.60, cache read $0.06. official $0.75 / $3.75 (vendor list price, checked 2026-10-02), save 20% / 4%. Chat, reasoning, adaptive thinking, image input, very long context. ~1,000,000 tokens (million-token long context). docs: https://nezhagate.com/en/docs/gemini-3.7-flash. - [Gemini 3.6 Flash](https://nezhagate.com/en/model/gemini-3.6-flash): Google. input $0.60 / output $3.60, cache read $0.06. official $0.75 / $3.75 (vendor list price, checked 2026-10-02), save 20% / 4%. Chat, reasoning, thinking, image input, very long context. ~1,000,000 tokens (million-token long context). docs: https://nezhagate.com/en/docs/gemini-3.6-flash. - [Gemini 3.6 Flash High](https://nezhagate.com/en/model/gemini-3.6-flash-high): Google. input $0.60 / output $3.60, cache read $0.06. official $0.75 / $3.75 (vendor list price, checked 2026-10-02), save 20% / 4%. Deep reasoning, complex tasks, thinking, image input, very long context. ~1,000,000 tokens (million-token long context). docs: https://nezhagate.com/en/docs/gemini-3.6-flash-high. - [Gemini 3.6 Flash Low](https://nezhagate.com/en/model/gemini-3.6-flash-low): Google. input $0.60 / output $3.60, cache read $0.06. official $0.75 / $3.75 (vendor list price, checked 2026-10-02), save 20% / 4%. High-speed chat, high volume, low latency, image input. ~1,000,000 tokens (million-token long context). docs: https://nezhagate.com/en/docs/gemini-3.6-flash-low. - [Gemini 3.6 Flash Tiered](https://nezhagate.com/en/model/gemini-3.6-flash-tiered): Google. input $0.60 / output $3.60, cache read $0.06. official $0.75 / $3.75 (vendor list price, checked 2026-10-02), save 20% / 4%. Adaptive thinking, chat, reasoning, image input. ~1,000,000 tokens (million-token long context). docs: https://nezhagate.com/en/docs/gemini-3.6-flash-tiered. - [Gemini 3 Flash](https://nezhagate.com/en/model/gemini-3-flash-preview): Google. input $0.30 / output $1.20, cache read $0.03. official $0.50 / $3.00 (vendor list price, checked 2026-10-02), save 40% / 60%. Chat, reasoning, high concurrency, low latency. ~1,000,000 tokens (million-token long context). docs: https://nezhagate.com/en/docs/gemini-3-flash-preview. - [Gemini 2.5 Flash](https://nezhagate.com/en/model/gemini-2.5-flash): Google. input $0.30 / output $1.20, cache read $0.03. official $0.30 / $2.50 (vendor list price, checked 2026-10-02), save 52% on output. Chat, high concurrency, low latency, multimodal. ~1,000,000 tokens (million-token long context). docs: https://nezhagate.com/en/docs/gemini-2.5-flash. - [DeepSeek V4.1 Flash](https://nezhagate.com/en/model/deepseek-v4.1-flash): DeepSeek. input $0.09 / output $0.36, cache read $0.002. official $0.15 / $0.60 (vendor list price, checked 2026-10-02), save 40%. Chat, reasoning, coding, tool calling, switchable thinking mode. measured on real customer calls over the last 7 days: 85.5% success. docs: https://nezhagate.com/en/docs/deepseek-v4.1-flash. - [DeepSeek V4 Flash 0731](https://nezhagate.com/en/model/deepseek-v4-flash-0731): DeepSeek. input $0.045 / output $0.18, cache read $0.0009. official $0.15 / $0.60 (vendor list price, checked 2026-10-02), save 70%. Chat, reasoning, coding, tool calling, switchable thinking mode, pinned version. docs: https://nezhagate.com/en/docs/deepseek-v4-flash-0731. - [GLM-5.3](https://nezhagate.com/en/model/glm-5.3): Z.ai. input $0.34 / output $1.25, cache read $0.09. official $1.40 / $4.40 (vendor list price, checked 2026-10-02), save 75% / 71%. Chat, coding, agentic workflows, reasoning, Chinese and English writing, tool calling. docs: https://nezhagate.com/en/docs/glm-5.3. - [GLM-5.3 Flash](https://nezhagate.com/en/model/glm-5.3-flash): Z.ai. input $0.072 / output $0.25, cache read $0.02. official $0.15 / $0.50 (vendor list price, checked 2026-10-02), save 52% / 50%. High-volume chat, classification and extraction, summarising and rewriting, tool calling. measured on real customer calls over the last 7 days: 80.2% success. docs: https://nezhagate.com/en/docs/glm-5.3-flash. - [Kimi K3](https://nezhagate.com/en/model/kimi-k3): Moonshot. input $2.10 / output $10.50, cache read $0.21. official $3.00 / $15.00 (vendor list price, checked 2026-10-02), save 30%. Long-document understanding, multi-document analysis, Chinese writing, agents and tool calling, reasoning. docs: https://nezhagate.com/en/docs/kimi-k3. - [Qwen3.7 Max](https://nezhagate.com/en/model/qwen3.7-max): Alibaba. input $1.65 / output $4.85, cache read $0.33. official $2.50 / $7.50 (vendor list price, checked 2026-10-02), save 34% / 35%. Reasoning, coding, agents and tool calling, Chinese and English writing, long context. docs: https://nezhagate.com/en/docs/qwen3.7-max. - [Qwen3.8 Max](https://nezhagate.com/en/model/qwen3.8-max): Alibaba. input $1.40 / output $4.20, cache read $0.175. official $2.00 / $6.00 (vendor list price, checked 2026-10-02), save 30%. Reasoning, coding, agents and tool calling, Chinese and English writing, image understanding. docs: https://nezhagate.com/en/docs/qwen3.8-max. - [Qwen3.8 Max 0902](https://nezhagate.com/en/model/qwen3.8-max-0902): Alibaba. input $1.40 / output $4.20, cache read $0.175. official $2.00 / $6.00 (vendor list price, checked 2026-10-02), save 30%. Reasoning, coding, agents and tool calling, image understanding, pinned version. docs: https://nezhagate.com/en/docs/qwen3.8-max-0902. - [Qwen3.8 Flash](https://nezhagate.com/en/model/qwen3.8-flash): Alibaba. input $0.10 / output $0.33, cache read $0.012. official $0.15 / $0.47 (vendor list price, checked 2026-10-02), save 33% / 29%. High-volume chat, classification and extraction, summarising and rewriting, image understanding and OCR, tool calling. measured on real customer calls over the last 7 days: 96.8% success. docs: https://nezhagate.com/en/docs/qwen3.8-flash. - [Doubao Seed 2.1 Pro](https://nezhagate.com/en/model/doubao-seed-2-1-pro): ByteDance. input $0.71 / output $3.57, cache read $0.14. official $0.89 / $4.46 (vendor list price, checked 2026-10-02), save 20% / 19%. Reasoning, coding, agents and tool calling, JSON output, image understanding. docs: https://nezhagate.com/en/docs/doubao-seed-2-1-pro. - [Doubao Seed 2.1 Turbo](https://nezhagate.com/en/model/doubao-seed-2-1-turbo): ByteDance. input $0.36 / output $1.79, cache read $0.07. official $0.45 / $2.23 (vendor list price, checked 2026-10-02), save 20% / 19%. Everyday chat, classification and extraction, summarising and rewriting, image understanding, tool calling. docs: https://nezhagate.com/en/docs/doubao-seed-2-1-turbo. - [Grok 4.7](https://nezhagate.com/en/model/grok-4.7): xAI. input $0.30 / output $0.90, cache read $0.075. official $2.00 / $6.00 (vendor list price, checked 2026-10-02), save 85%. Complex reasoning, coding, long-document analysis, tool calling, structured output, image understanding. docs: https://nezhagate.com/en/docs/grok-4.7. ## Image models (live prices, USD per image) - [GPT Image 2](https://nezhagate.com/en/model/gpt-image-2): OpenAI. $0.005 (1K) / $0.01 (2K) / $0.02 (4K) per image. text-to-image and image editing (/v1/images/edits). measured on real customer calls over the last 7 days: 99.8% success, median 31 s per image (p90 42 s). providers compared: https://nezhagate.com/en/compare/gpt-image. docs: https://nezhagate.com/en/docs/gpt-image-2. - [GPT Image 2.5 Flare](https://nezhagate.com/en/model/gpt-image-2.5-flare): OpenAI. $0.005 (1K) / $0.01 (2K) / $0.02 (4K) per image. text-to-image and image editing (/v1/images/edits). measured on real customer calls over the last 7 days: 100% success, median 24 s per image (p90 38 s). providers compared: https://nezhagate.com/en/compare/gpt-image. docs: https://nezhagate.com/en/docs/gpt-image-2.5-flare. - [GPT Image 2.5 Sunburst](https://nezhagate.com/en/model/gpt-image-2.5-sunburst): OpenAI. $0.005 (1K) / $0.01 (2K) / $0.02 (4K) per image. text-to-image and image editing (/v1/images/edits). measured on real customer calls over the last 7 days: 99.9% success, median 30 s per image (p90 47 s). docs: https://nezhagate.com/en/docs/gpt-image-2.5-sunburst. - [Nano Banana 2](https://nezhagate.com/en/model/nano-banana-2): Google. $0.015 (1K) / $0.025 (2K) / $0.04 (4K) per image. official $0.067 / $0.101 / $0.151 (vendor list price, checked 2026-10-02), save 77% / 75% / 73%. text-to-image and image editing (/v1/images/edits). measured on real customer calls over the last 7 days: 100% success, median 18 s per image (p90 32 s). providers compared: https://nezhagate.com/en/compare/nano-banana. docs: https://nezhagate.com/en/docs/nano-banana-2. - [Nano Banana Pro](https://nezhagate.com/en/model/nano-banana-pro): Google. $0.02 (1K) / $0.03 (2K) / $0.05 (4K) per image. official $0.134 / $0.134 / $0.24 (vendor list price, checked 2026-10-02), save 85% / 77% / 79%. text-to-image and image editing (/v1/images/edits). providers compared: https://nezhagate.com/en/compare/nano-banana. docs: https://nezhagate.com/en/docs/nano-banana-pro. ## Video models (live prices, USD) - [Veo 3.1](https://nezhagate.com/en/model/veo-3.1): Google. 720p $0.20 per clip, 1080p $0.22 per clip. fixed 8 s per clip. up to 3 reference images. first frame, first + last frame, or subject references. native audio always on. providers compared: https://nezhagate.com/en/compare/veo. docs: https://nezhagate.com/en/docs/veo-3.1. - [Gemini Omni Flash](https://nezhagate.com/en/model/gemini-omni-flash): Google. 720P $0.05/s, 1080P $0.05/s. 6 / 8 / 10 s per clip. up to 3 reference images. first frame, first + last frame, or subject references. prompt up to 20,000 characters. native audio always on. providers compared: https://nezhagate.com/en/compare/veo. docs: https://nezhagate.com/en/docs/gemini-omni-flash. - [Seedance 2.5](https://nezhagate.com/en/model/seedance-2.5): ByteDance. 480P $0.10/s, 720P $0.158/s (1080P coming soon). official 480P $0.103/s, 720P $0.231/s (vendor list price, checked 2026-10-02), save 480P 2%, 720P 31%. 4-30 s per clip. up to 30 reference images. 10 reference videos (30 s in total). 10 reference audio clips. images are references (@Image1, @Image2 ...), no first-frame control. prompt up to 11,000 characters. native audio, can be switched off. reference videos are not billed. measured on real customer calls over the last 7 days: 86.5% success, median 5.8 min per clip (p90 9.8 min). providers compared: https://nezhagate.com/en/compare/seedance. docs: https://nezhagate.com/en/docs/seedance-2.5. - [Seedance 2.0](https://nezhagate.com/en/model/seedance-2.0): ByteDance. 480P $0.07/s, 720P $0.15/s (1080P coming soon). 5-15 s per clip. up to 9 reference images. 3 reference videos (15 s in total). 3 reference audio clips. images are references (@Image1, @Image2 ...), no first-frame control. native audio, can be switched off. reference videos are not billed. providers compared: https://nezhagate.com/en/compare/seedance. docs: https://nezhagate.com/en/docs/seedance-2.0. - [Seedance 2.0 Fast](https://nezhagate.com/en/model/seedance-2.0-fast): ByteDance. 720P $0.075/s. official 720P $0.121/s (vendor list price, checked 2026-10-02), save 720P 38%. 5-15 s per clip. up to 9 reference images. 3 reference videos (15 s in total). 3 reference audio clips. images are references (@Image1, @Image2 ...), no first-frame control. native audio, can be switched off. reference videos are not billed. providers compared: https://nezhagate.com/en/compare/seedance. docs: https://nezhagate.com/en/docs/seedance-2.0-fast. - [Seedance 2.5 · 30s](https://nezhagate.com/en/model/seedance-2.5-30s): ByteDance. 720P $1.50 per clip. official $6.93 per clip (vendor list price, checked 2026-10-02), save 78%. fixed 30 s per clip. up to 9 reference images. images are references (@Image1, @Image2 ...), no first-frame control. native audio always on. measured on real customer calls over the last 7 days: 91.9% success, median 12.8 min per clip (p90 20.9 min). providers compared: https://nezhagate.com/en/compare/seedance. docs: https://nezhagate.com/en/docs/seedance-2.5-30s. - [Wan 3.0](https://nezhagate.com/en/model/wan3.0-video): Alibaba. 480P $0.04/s, 720P $0.08/s, 1080P $0.16/s. official 480P $0.05/s, 720P $0.10/s, 1080P $0.20/s (vendor list price, checked 2026-10-02), save 480P 20%, 720P 20%, 1080P 20%. 2-30 s per clip. up to 10 reference images. 5 reference videos (15 s in total). first frame, first + last frame, or subject references. no audio switch. reference-video seconds are billed. providers compared: https://nezhagate.com/en/compare/wan. docs: https://nezhagate.com/en/docs/wan3.0-video. - [Wan 3.0 Prime](https://nezhagate.com/en/model/wan3.0-video-prime): Alibaba. 480P $0.054/s, 720P $0.112/s, 1080P $0.20/s. official 480P $0.068/s, 720P $0.14/s, 1080P $0.28/s (vendor list price, checked 2026-10-02), save 480P 20%, 720P 20%, 1080P 28%. 2-30 s per clip. up to 10 reference images. 5 reference videos (15 s in total). first frame, first + last frame, or subject references. no audio switch. reference-video seconds are billed. providers compared: https://nezhagate.com/en/compare/wan. docs: https://nezhagate.com/en/docs/wan3.0-video-prime. - [MiniMax H3](https://nezhagate.com/en/model/minimax-h3): MiniMax. 1080P $0.036/s. official 768P $0.08/s, 2K $0.13/s (vendor list price, checked 2026-10-02; no 1080P tier), save 55% against 768P. 5-15 s per clip. up to 5 reference images. 3 reference audio clips. first-frame mode by default; image_role=reference for multi-image reference. providers compared: https://nezhagate.com/en/compare/minimax-h3. docs: https://nezhagate.com/en/docs/minimax-h3. - [Grok Imagine Video 1.5](https://nezhagate.com/en/model/grok-imagine-video-1.5): xAI. 480P $0.30 per clip, 720P $0.40 per clip. 4-15 s per clip. up to 7 reference images. native audio always on. docs: https://nezhagate.com/en/docs/grok-imagine-video-1.5. ## Coming soon (not callable yet) - [Claude Opus 5.5](https://nezhagate.com/en/model/claude-opus-5-5): coming soon - [Gemini 4 Argon](https://nezhagate.com/en/model/gemini-4-argon): coming soon; Gemini 4 Argon API: Google's first Gemini 4 frontier model, built for deep reasoning over long workflows, with a 1M-token output limit - [Claude Sonnet 5.5](https://nezhagate.com/en/model/claude-sonnet-5-5): coming soon; Claude Sonnet 5.5 API: Anthropic's newest Sonnet, the best combination of speed and intelligence, 30%+ faster than Sonnet 5 at the same price - [Claude Fable 5.1](https://nezhagate.com/en/model/claude-fable-5-1): coming soon; Anthropic's most capable widely released model: the successor to Fable 5, with a 1M-token context ## Price comparisons - [Seedance API Pricing Compared: 2.5, 2.0, Fast](https://nezhagate.com/en/compare/seedance): official list price, OpenRouter, fal, Kie, Replicate, APIMart and NezhaGate side by side (checked 2026-10-03) - [Nano Banana 2 & Pro API Pricing Compared](https://nezhagate.com/en/compare/nano-banana): official list price, OpenRouter, fal, Kie, Replicate, APIMart and NezhaGate side by side (checked 2026-10-03) - [GPT Image 2 & 2.5 API Pricing Compared](https://nezhagate.com/en/compare/gpt-image): official list price, OpenRouter, fal, Kie, Replicate, APIMart and NezhaGate side by side (checked 2026-10-03) - [MiniMax H3 (Hailuo 3) API Pricing Compared](https://nezhagate.com/en/compare/minimax-h3): official list price, OpenRouter, fal, Kie, Replicate, APIMart and NezhaGate side by side (checked 2026-10-03) - [Veo 3.1 & Gemini Omni Flash API Pricing Compared](https://nezhagate.com/en/compare/veo): official list price, OpenRouter, fal, Kie, Replicate, APIMart and NezhaGate side by side (checked 2026-10-03) - [Wan 3.0 & Wan 3.0 Prime API Pricing Compared](https://nezhagate.com/en/compare/wan): official list price, OpenRouter, fal, Kie, Replicate, APIMart and NezhaGate side by side (checked 2026-10-03) ## Guides - [Seedance 2.5 API: Price, Parameters and Code](https://nezhagate.com/en/learn/seedance-2-5-api): Call the Seedance 2.5 video API: $0.158 per second at 720P, 4 to 30 seconds, up to 30 reference images, 10 videos and 10 audio clips, with copy-paste curl, Python and Node.js code. - [Seedance 30-Second Video API: Price, Samples, Code](https://nezhagate.com/en/learn/seedance-30-second-video): Seedance 2.5 renders 30 seconds in one shot. Two ways on NezhaGate: seedance-2.5-30s at $1.50 per 720P clip, or seedance-2.5 per second. Samples with full prompts, prompt structure and code. ## Docs and machine-readable files - [API docs and integration guide](https://nezhagate.com/en/docs-guide) - [Pricing JSON: machine-readable sell prices (New API compatible)](https://nezhagate.com/api/pricing) - [Agent skill for Claude Code / Codex: npx skills add https://nezhagate.com](https://nezhagate.com/en/docs-guide#skills) - [Client setup guides (SillyTavern, Cherry Studio, Cline)](https://nezhagate.com/en/guides) ## Price table | Model ID | Name | Type | NezhaGate price | Official (in / out) | Vendor | | --- | --- | --- | --- | --- | --- | | gpt-5.6-sol | GPT-5.6 Sol | chat / reasoning | $0.40 / $2.00 | $4.00 / $20.00 | OpenAI | | gpt-5.6-terra | GPT-5.6 Terra | chat / reasoning | $0.20 / $1.20 | $2.00 / $12.00 | OpenAI | | gpt-5.6-luna | GPT-5.6 Luna | chat / reasoning | $0.07 / $0.40 | $0.20 / $1.20 | OpenAI | | gpt-5.5 | GPT-5.5 | chat / reasoning | $0.50 / $3.00 | $5.00 / $30.00 | OpenAI | | gpt-6-astra | GPT-6 Astra | chat / reasoning | $1.00 / $5.00 | $10.00 / $50.00 | OpenAI | | gpt-6.1-sol | GPT-6.1 Sol | chat / reasoning | $0.40 / $2.00 | $2.00 / $10.00 | OpenAI | | gpt-6-sol | GPT-6 Sol | chat / reasoning | $0.20 / $1.00 | $2.00 / $10.00 | OpenAI | | gpt-6-luna | GPT-6 Luna | chat / reasoning | $0.07 / $0.40 | $0.10 / $0.50 | OpenAI | | gpt-image-2 | GPT Image 2 | image generation | $0.005 / $0.01 / $0.02 per image | - | OpenAI | | gpt-image-2.5-flare | GPT Image 2.5 Flare | image generation | $0.005 / $0.01 / $0.02 per image | - | OpenAI | | gpt-image-2.5-sunburst | GPT Image 2.5 Sunburst | image generation | $0.005 / $0.01 / $0.02 per image | - | OpenAI | | nano-banana-2 | Nano Banana 2 | image generation | $0.015 / $0.025 / $0.04 per image | - | Google | | nano-banana-pro | Nano Banana Pro | image generation | $0.02 / $0.03 / $0.05 per image | - | Google | | claude-sonnet-4-6 | Claude Sonnet 4.6 | chat / reasoning | $1.50 / $7.50 | $3.00 / $15.00 | Anthropic | | claude-opus-5 | Claude Opus 5 | chat / reasoning | $4.00 / $20.00 | $5.00 / $25.00 | Anthropic | | claude-fable-5 | Claude Fable 5 | chat / reasoning | $8.00 / $40.00 | $10.00 / $50.00 | Anthropic | | claude-sonnet-5 | Claude Sonnet 5 | chat / reasoning | $1.00 / $5.00 | $2.00 / $10.00 | Anthropic | | claude-opus-5-5 | Claude Opus 5.5 | chat / reasoning | coming soon | $4.00 / $20.00 | Anthropic | | gemini-3.1-pro | Gemini 3.1 Pro | chat / reasoning | $0.50 / $3.00 | $2.00 / $12.00 | Google | | gemini-3.8-flash | Gemini 3.8 Flash | chat / reasoning | $0.60 / $3.60 | $0.75 / $3.75 | Google | | gemini-3.7-flash | Gemini 3.7 Flash | chat / reasoning | $0.60 / $3.60 | $0.75 / $3.75 | Google | | gemini-3.6-flash | Gemini 3.6 Flash | chat / reasoning | $0.60 / $3.60 | $0.75 / $3.75 | Google | | gemini-3.6-flash-high | Gemini 3.6 Flash High | chat / reasoning | $0.60 / $3.60 | $0.75 / $3.75 | Google | | gemini-3.6-flash-low | Gemini 3.6 Flash Low | chat / reasoning | $0.60 / $3.60 | $0.75 / $3.75 | Google | | gemini-3.6-flash-tiered | Gemini 3.6 Flash Tiered | chat / reasoning | $0.60 / $3.60 | $0.75 / $3.75 | Google | | gemini-3-flash-preview | Gemini 3 Flash | chat / reasoning | $0.30 / $1.20 | $0.50 / $3.00 | Google | | gemini-2.5-flash | Gemini 2.5 Flash | chat / reasoning | $0.30 / $1.20 | $0.30 / $2.50 | Google | | deepseek-v4.1-flash | DeepSeek V4.1 Flash | chat / reasoning | $0.09 / $0.36 | $0.15 / $0.60 | DeepSeek | | deepseek-v4-flash-0731 | DeepSeek V4 Flash 0731 | chat / reasoning | $0.045 / $0.18 | $0.15 / $0.60 | DeepSeek | | glm-5.3 | GLM-5.3 | chat / reasoning | $0.34 / $1.25 | $1.40 / $4.40 | Z.ai | | glm-5.3-flash | GLM-5.3 Flash | chat / reasoning | $0.072 / $0.25 | $0.15 / $0.50 | Z.ai | | kimi-k3 | Kimi K3 | chat / reasoning | $2.10 / $10.50 | $3.00 / $15.00 | Moonshot | | qwen3.7-max | Qwen3.7 Max | chat / reasoning | $1.65 / $4.85 | $2.50 / $7.50 | Alibaba | | qwen3.8-max | Qwen3.8 Max | chat / reasoning | $1.40 / $4.20 | $2.00 / $6.00 | Alibaba | | qwen3.8-max-0902 | Qwen3.8 Max 0902 | chat / reasoning | $1.40 / $4.20 | $2.00 / $6.00 | Alibaba | | qwen3.8-flash | Qwen3.8 Flash | chat / reasoning | $0.10 / $0.33 | $0.15 / $0.47 | Alibaba | | doubao-seed-2-1-pro | Doubao Seed 2.1 Pro | chat / reasoning | $0.71 / $3.57 | $0.89 / $4.46 | ByteDance | | doubao-seed-2-1-turbo | Doubao Seed 2.1 Turbo | chat / reasoning | $0.36 / $1.79 | $0.45 / $2.23 | ByteDance | | grok-4.7 | Grok 4.7 | chat / reasoning | $0.30 / $0.90 | $2.00 / $6.00 | xAI | | veo-3.1 | Veo 3.1 | video generation | 720p $0.20 per clip, 1080p $0.22 per clip | - | Google | | gemini-omni-flash | Gemini Omni Flash | video generation | 720P $0.05/s, 1080P $0.05/s | - | Google | | seedance-2.5 | Seedance 2.5 | video generation | 480P $0.10/s, 720P $0.158/s | - | ByteDance | | seedance-2.0 | Seedance 2.0 | video generation | 480P $0.07/s, 720P $0.15/s | - | ByteDance | | seedance-2.0-fast | Seedance 2.0 Fast | video generation | 720P $0.075/s | - | ByteDance | | seedance-2.5-30s | Seedance 2.5 · 30s | video generation | 720P $1.50 per clip | - | ByteDance | | wan3.0-video | Wan 3.0 | video generation | 480P $0.04/s, 720P $0.08/s, 1080P $0.16/s | - | Alibaba | | wan3.0-video-prime | Wan 3.0 Prime | video generation | 480P $0.054/s, 720P $0.112/s, 1080P $0.20/s | - | Alibaba | | minimax-h3 | MiniMax H3 | video generation | 1080P $0.036/s | - | MiniMax | | grok-imagine-video-1.5 | Grok Imagine Video 1.5 | video generation | 480P $0.30 per clip, 720P $0.40 per clip | - | xAI | ## Code samples ### Chat (OpenAI SDK) ```python from openai import OpenAI client = OpenAI( base_url="https://nezhagate.com/v1", api_key="YOUR_NEZHAGATE_API_KEY", ) stream = client.chat.completions.create( model="gpt-6-astra", messages=[{"role": "user", "content": "Explain rate limiting in two sentences."}], stream=True, ) for chunk in stream: if chunk.choices and chunk.choices[0].delta.content: print(chunk.choices[0].delta.content, end="", flush=True) ``` ### curl ```bash curl https://nezhagate.com/v1/chat/completions \ -H "Authorization: Bearer $NEZHAGATE_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "gpt-6-astra", "messages": [{"role": "user", "content": "Explain rate limiting in two sentences."}], "stream": true }' ``` ### Image (async job) ```bash # 1) submit: returns HTTP 202 and a job id curl https://nezhagate.com/v1/images/generations \ -H "Authorization: Bearer $NEZHAGATE_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model": "gpt-image-2", "prompt": "A red fox in fresh snow, soft morning light", "size": "16:9", "resolution": "1K"}' # 2) poll every 3 s until "status" is "succeeded", then read data[0].url curl https://nezhagate.com/v1/images/jobs/JOB_ID \ -H "Authorization: Bearer $NEZHAGATE_API_KEY" ``` ### Video (async job) ```python import time import requests BASE = "https://nezhagate.com/v1" HEADERS = {"Authorization": "Bearer YOUR_NEZHAGATE_API_KEY"} job = requests.post(f"{BASE}/videos/generations", headers=HEADERS, json={ "model": "seedance-2.5", "prompt": "A paper boat drifting down a rainy street at night, neon reflections", "seconds": 5, "resolution": "720P", "size": "16:9", }).json() while True: r = requests.get(f"{BASE}/videos/jobs/{job['id']}", headers=HEADERS).json() if r["status"] == "succeeded": print(r["data"][0]["url"]) break if r["status"] == "failed": raise RuntimeError(r["error"]["message"]) time.sleep(15) ``` ## FAQ ### What is NezhaGate? NezhaGate is an OpenAI-compatible AI API platform: one key gives you chat, image and video models. Existing OpenAI SDK code connects by changing the base URL to https://nezhagate.com/v1 and using your NezhaGate key, and Claude models also speak the native Anthropic API. ### What does "chat, images, video — one API" actually mean? One Base URL and one API key cover all three kinds of work: chat at /v1/chat/completions, images at /v1/images/generations (image-to-image at /v1/images/edits), and video at /v1/videos/generations. The models come from OpenAI (GPT), Anthropic (Claude), Google (Gemini), ByteDance (Seedance), Alibaba (Wan), xAI (Grok), MiniMax and more, so you do not integrate an SDK, open an account and reconcile a bill for each vendor separately. To switch or compare models, you change a single model parameter. ### How does NezhaGate differ from the official OpenAI / Claude / Gemini APIs? NezhaGate serves the vendors' own models; what changes is how you reach them. One key covers every vendor, the API is OpenAI-compatible across the board (Claude also speaks the native Anthropic API), and billing, logs and balance live in one console. Flagship models are connected to multiple suppliers with automatic failover, and you save up to 90% against official prices. ### Which models are supported? NezhaGate currently offers OpenAI's GPT-6 family (Astra / 6.1 Sol / Sol / Luna), the GPT-5.6 family (Sol / Terra / Luna), GPT-5.5 and the image models GPT Image 2 and GPT Image 2.5 (Flare / Sunburst); Anthropic's Claude Opus 5, Claude Sonnet 5, Claude Sonnet 4.6 and Claude Fable 5; Google's Gemini 3.1 Pro, Gemini 3.8 / 3.7 / 3.6 / 3 / 2.5 Flash and the image models Nano Banana 2 and Nano Banana Pro; DeepSeek V4.1 Flash and V4 Flash 0731; Z.ai's GLM-5.3 and GLM-5.3 Flash; Moonshot's Kimi K3; Alibaba's Qwen3.8 family (Max / Max 0902 / Flash) and Qwen3.7 Max; ByteDance's Doubao Seed 2.1 (Pro / Turbo); xAI's Grok 4.7; and for video, Seedance 2.5 and the Seedance 2.0 family, Wan 3.0, MiniMax H3 and Grok Imagine Video 1.5. Claude Opus 5.5 and Grok 4.6 are coming soon. Every model shares the same endpoint and key; the live, complete list is on the Model Marketplace. ### Are NezhaGate's models the original vendor models? Yes. Every answer on NezhaGate comes from the model you asked for: we never swap in another model and never rewrite its output. Every call is in your logs with the model, usage, charge, latency and whether it switched lines, and each model's live state is public on the Status page. Every week we also run a fixed question set against our flagship models to measure success rate, time to first token and output speed. ### Do you support GPT, Claude and Gemini? Yes. The main models from GPT (OpenAI), Claude (Anthropic) and Gemini (Google) are all available through one OpenAI-compatible interface. Switching models only requires changing the model field in your request — no new SDK or re-integration needed. ### Do you support image and video generation? Yes. For images there are GPT Image 2, GPT Image 2.5 (Flare / Sunburst), Nano Banana 2 and Nano Banana Pro: text-to-image at /v1/images/generations and image-to-image at /v1/images/edits, up to 4K. For video there are Seedance, Wan 3.0, MiniMax H3 and Grok Imagine Video at /v1/videos/generations. Results come back as direct links you can use in your app right away. ### How often are prices and model details updated? The Pricing page and the Model Marketplace are wired to the billing system in real time, so the price you see is exactly the price you are charged. New models and price changes appear as soon as they go live, and price changes are announced on the site one day in advance. ### How is NezhaGate priced? Are there hidden fees? NezhaGate bills by actual usage: chat by the token, images by the picture, and video mostly by the second (a few models are priced per clip, as the Pricing page shows). You save up to 90% against official prices, and every model is listed next to its official price on the Pricing page — no hidden multipliers, no prices that only appear after you top up. Price changes are announced on the site one day in advance. ### Why is NezhaGate cheaper than the official APIs? NezhaGate buys model capacity in bulk from multiple suppliers at below-retail cost, so most models cost less here than at the official API; the Pricing page shows exactly how much you save on each one. We earn the margin between our cost and our price, and we charge no top-up fees. ### How are failed calls billed? NezhaGate never charges for a failed call. Asynchronous jobs such as images and video reserve credits when you submit them and refund them in full automatically if the job fails, with the reason written to your log. If a streamed answer breaks off midway, you are billed only for the usage the upstream actually returned. ### Do NezhaGate credits expire? Can I get a refund? NezhaGate credits don't expire. Top-ups start at ¥10 (260 credits) with WeChat Pay and at $5 with USDT or a card, and every new account gets 80 free credits at sign-up. Your top-up history and the charge for every call in the last 60 days are in the console. Within 7 days of a top-up, unused purchased credits can be refunded through support, minus a $0.40 fee per refund; bonus credits are not refundable, and USDT payments can be refunded too. ### How do I check my balance and usage? Sign in and the console shows your live balance, usage per API key, and every call from the last 60 days: model, time, tokens or image count, latency and charge, with the reason written next to any failed call. You can also read your balance programmatically through the balance endpoint. ### How reliable is NezhaGate? NezhaGate connects flagship models — GPT, Gemini, DeepSeek, GLM, Kimi, GPT Image, Nano Banana, Seedance 2.0 and more — to multiple suppliers. If one returns an error, times out or is at capacity, the request moves automatically to the next supplier for the same model; a streamed answer switches before it starts and is never moved once output has begun. The gateway is probed around the clock, every 2 minutes; live model status and gateway availability are published on the Status page, with success rates counted from real customer traffic only. Failed calls are never billed, and failed image and video jobs are refunded in full. ### Does NezhaGate have rate limits? NezhaGate sets no default per-account rate limit. On the API Keys page you can give each key its own requests-per-minute cap, daily budget, total spend limit, model allowlist and IP allowlist. If every line for a model is at capacity at once, you get a 429 with a Retry-After header; retry after that interval. For sustained high concurrency, contact support in advance and we will add capacity for your volume. Free promotional models carry a daily call quota. ### Is it suitable for commercial projects or AI tool sites? Yes. NezhaGate gives you an OpenAI-compatible API, per-project API keys (each with its own spend, rate, model and IP limits), itemised call logs and billing that never charges for failures, ready to plug into commercial products and tool sites. Please follow the Terms of Service and each model vendor's usage policies. ### Do you support enterprise cooperation or bulk calls? Yes. NezhaGate can handle larger-scale batch calls, and we welcome enterprise inquiries. For higher concurrency, dedicated quota or integration needs, email support@nezhagate.com or use the Contact us option on the site. The plan is tailored to your usage and scenario. ### Which scenarios suit NezhaGate best? NezhaGate suits cases where you need to call several model providers in one project, want unified billing and usage management, or want to quickly validate different models. Common uses include AI tool sites, SaaS products, automation workflows, agents and multi-model comparison. It is especially convenient when you need to switch flexibly between models. ### How do I get started quickly with NezhaGate? Getting started takes three steps: register an account, create an API key in the console, then point your Base URL at NezhaGate and add the key. Because it is fully OpenAI-compatible, existing OpenAI SDK code needs almost no changes — just swap base_url and api_key. See the developer docs for details. ### How should I configure the Base URL? Set the Base URL to https://nezhagate.com/v1. In the OpenAI SDK, set base_url (or the OPENAI_BASE_URL environment variable) to that address and use the key you created in the console. Use /v1/chat/completions for chat and /v1/images/generations for images — the paths match OpenAI. ### Is it compatible with the OpenAI SDK? Yes. NezhaGate implements the OpenAI-compatible interface, so the official openai SDK works by simply changing base_url and api_key. Request and response shapes match OpenAI, including streaming output and multimodal input. ### Do you support Python, Node.js and curl? Yes. Any language that can make HTTP requests or use the OpenAI SDK works, including Python, Node.js, Go, Java and plain curl. The docs include ready-to-run Python, Node.js and curl examples. ### Do you support n8n, Dify, LangChain, Open WebUI, LobeChat, Cherry Studio and ChatBox? Yes. Most of these tools let you set a custom OpenAI-compatible Base URL and API key — point them at NezhaGate to use the platform's models. If a tool supports an OpenAI-compatible or custom endpoint, it can connect to NezhaGate with no extra plugin. ### What if my API key is leaked? If you suspect an API key is leaked, delete or rotate it in the console immediately — once revoked, the old key can no longer be used. We recommend creating separate keys per project and never embedding keys in front-end code or public repos. You can check each key's usage in the console at any time to spot anomalies. ### Does NezhaGate store or look at my data? NezhaGate uses call content only to complete the call, to show it to you in your logs and to troubleshoot problems; it is never used to train any model and never sold or shared with advertisers. Apart from the model suppliers and infrastructure services needed to complete a call (CDN, cloud storage, and a search service when you turn on web search), it is not shared with third parties. Call logs stay in your console for 60 days; their text content is deleted after 30 days, leaving only the usage record needed for billing. Generated images and videos are deleted after 60 days and uploaded reference files after 7 days. Database backups are kept for up to 30 more days. ### Are my account and API keys safe? Every API key can carry its own spend limit, requests-per-minute cap, model allowlist and IP allowlist, and can be disabled or deleted at any time, so one leaked key never reaches your other projects. Keys are stored encrypted on our servers and you can view and copy them whenever you are signed in; every call and every balance change is itemised for you to check. ### Who do I contact if something goes wrong? NezhaGate support is under "Contact" in the site's top navigation and in your console: WeChat, our Telegram group, and email at support@nezhagate.com. Support hours are 10:00–19:00 Beijing time (UTC+8), and we reply within an hour of seeing your message. ### What uses are prohibited? You may not use NezhaGate for anything that violates laws or the model vendors' usage policies, including generating or distributing illegal, infringing, fraudulent or malware content, or automated abuse and bypassing quota or safety limits. Violations may lead to account restriction or termination. The full rules are in the Terms of Service. ### What should I keep in mind when using NezhaGate? Keep your API keys safe, top up as you need, and follow the Terms of Service and each model vendor's usage policies. Check unit prices on the Pricing page before you call, and give your keys a daily budget and a total limit to rule out unexpected spend. ## Optional - https://nezhagate.com/en/llms.txt - https://nezhagate.com/en/pricing - https://nezhagate.com/en/docs-guide - https://nezhagate.com/en/faq