NezhaGate
Chat

GPT-6 Astra

gpt-6-astra
24H STATUS Success: 100.0%
-24hnow
Only successes and our own failures are counted: green is success in that hour, red is the failure share. Two things are excluded -- client-side errors (bad request, auth, insufficient balance and other 4xx) and requests still in flight. One bar per hour, taken from real gateway call logs.

GPT-6 Astra is the next-generation flagship OpenAI released in September 2026: a 1,050,000-token context window, up to 128,000 output tokens, reasoning_effort up to xhigh, image input, tool calling and prompt caching. Fully OpenAI-SDK compatible - set model to gpt-6-astra - and served on both /chat/completions and /responses. Pay-as-you-go, failed calls never billed; cached input bills at one tenth, and we do not add the long-context surcharge the official API applies.

ChatNext-gen flagship1.05M contextPrompt cachingImage input

Live Test · Playground

Try out GPT-6 Astra right here (available after login)。

Input

Advanced

This model is served over a ChatGPT-subscription upstream that does not accept temperature / top_p / max_tokens (they are silently ignored). Use reasoning_effort below for thinking depth, and prompt wording for length.

Web search
Memory · multi-turn

Conversation

Start a conversationType a message below to begin

About GPT-6 Astra

GPT-6 Astra is the next-generation flagship OpenAI released in September 2026, served on NezhaGate through the OpenAI-compatible API: a 1,050,000-token context window, up to 128,000 output tokens, reasoning_effort for thinking depth (up to xhigh), image input, tool calling and prompt caching, on both /chat/completions and /responses. Set model to gpt-6-astra and keep your SDK and request shape. Pay-as-you-go, failed calls never billed, and cached input settles at one tenth of the input rate.

Use cases

Very long context analysis

1,050,000 tokens holds a whole repository, a contract set or a long report in one call, with no chunking and stitching of your own.

Agentic coding

It speaks the OpenAI-compatible endpoint, so Codex, Cursor and Claude Code switch to GPT-6 Astra by changing the model field alone.

Deep reasoning work

reasoning_effort goes up to xhigh, so you can spend the thinking budget on the hardest part and keep simple requests on low.

Mixed image and text input

Image input with text output: screenshot understanding, chart questions and document parsing all use the same endpoint.

How to choose

Pick GPT-6 Astra when you need the strongest reasoning or a very long context; gpt-5.6-terra is cheaper for everyday development and gpt-5.6-luna for high-volume low-latency work. All three use the OpenAI-compatible endpoint, so switching is a one-field change.

FAQ

What are the context and output limits?
A 1,050,000-token context window and up to 128,000 output tokens per call. System prompts and other overhead count toward the context window.
How do I get cache hits?
Prompt caching is automatic: the cacheable prefix must be at least 1,024 tokens and byte-identical to the start of your previous request. Put the fixed parts - system prompt, tool definitions, long documents - at the very front and the changing user input last, and the hit rate is highest. The cached portion settles at one tenth of the input rate ($0.28 / 1M) and the console call log shows the cache-read token count.
How does the price compare with the official API?
The official list is $10 per 1M input and $50 per 1M output, and it doubles input and adds 50% to output once a prompt passes 272K tokens. NezhaGate is a flat $2.80 / $14.00 with no long-context surcharge, and cached reads at $0.28.
Is reasoning_effort supported?
Yes: low / medium / high / xhigh. More thinking means more output tokens, billed on actual usage; omit it and the model decides for itself.
Can I use /v1/responses?
Yes. The same model name works on /v1/chat/completions and /v1/responses, and both endpoints support streaming.

Related Models

Explore other models you can integrate.

View all →
GPT-5.6 Sol Chat

GPT-5.6 Sol

gpt-5.6-sol

Frontier flagship for agentic coding

$2.00/1M in · $12.00/1M out ↓75% View →
GPT-5.6 Terra Chat

GPT-5.6 Terra

gpt-5.6-terra

Balanced everyday agentic coding model

$1.20/1M in · $7.00/1M out ↓77% View →
GPT-5.6 Luna Chat

GPT-5.6 Luna

gpt-5.6-luna

Fast, economical agentic coding model

$0.80/1M in · $4.80/1M out ↓68% View →
GPT-5.5 Chat

GPT-5.5

gpt-5.5

Flagship chat and reasoning model

$0.70/1M in · $4.20/1M out ↓86% View →
GPT Image 2 Image

GPT Image 2

gpt-image-2

High-quality text-to-image / image-to-image model

$0.015/img and up View →
GPT Image 2.5 Flare Image

GPT Image 2.5 Flare

gpt-image-2.5-flare

Next-gen image model - clean and smooth

$0.015/img and up View →
GPT Image 2.5 Sunburst Image

GPT Image 2.5 Sunburst

gpt-image-2.5-sunburst

Next-gen image model - richer texture

$0.015/img and up View →
Nano Banana 2 Image

Nano Banana 2

nano-banana-2

High-quality text-to-image / image-to-image model

$0.025/img and up View →
Nano Banana Pro Image

Nano Banana Pro

nano-banana-pro

Flagship text-to-image / image-to-image model

$0.040/img and up View →
Claude Sonnet 4.6 Chat

Claude Sonnet 4.6

claude-sonnet-4-6

Balanced, efficient chat and coding model

$1.50/1M in · $7.50/1M out ↓50% View →
Claude Opus 5 Chat

Claude Opus 5

claude-opus-5

Next-generation flagship from Anthropic

$4.00/1M in · $20.00/1M out View →
Claude Fable 5 Chat

Claude Fable 5

claude-fable-5

Anthropic Fable series · narrative and long-form writing

$8.00/1M in · $40.00/1M out View →
Gemini 3.1 Pro Chat

Gemini 3.1 Pro

gemini-3.1-pro

Next-generation flagship reasoning model from Google

$0.50/1M in · $3.00/1M out ↓75% View →
Gemini 3.8 Flash Chat

Gemini 3.8 Flash

gemini-3.8-flash

Latest Flash · adaptive thinking

$0.60/1M in · $3.60/1M out ↓7% View →
Gemini 3.7 Flash Chat

Gemini 3.7 Flash

gemini-3.7-flash

Previous Flash · adaptive thinking

$0.60/1M in · $3.60/1M out ↓7% View →
Gemini 3.6 Flash Chat

Gemini 3.6 Flash

gemini-3.6-flash

Next-generation fast thinking model · four thinking budgets

$0.60/1M in · $3.60/1M out View →
Gemini 3.6 Flash High Chat

Gemini 3.6 Flash High

gemini-3.6-flash-high

Deep thinking tier

$0.60/1M in · $3.60/1M out View →
Gemini 3.6 Flash Low Chat

Gemini 3.6 Flash Low

gemini-3.6-flash-low

Fastest, lowest-cost tier

$0.60/1M in · $3.60/1M out View →
Gemini 3.6 Flash Tiered Chat

Gemini 3.6 Flash Tiered

gemini-3.6-flash-tiered

Auto-tiered thinking

$0.60/1M in · $3.60/1M out View →
Gemini 3 Flash Chat

Gemini 3 Flash

gemini-3-flash-preview

Fast, low-latency, cost-efficient chat model

$0.30/1M in · $1.20/1M out ↓57% View →
Gemini 3.5 Flash Chat

Gemini 3.5 Flash

gemini-3.5-flash

Next-generation fast thinking model

$0.45/1M in · $2.70/1M out ↓70% View →
Gemini 2.5 Flash Chat

Gemini 2.5 Flash

gemini-2.5-flash

Fast, low-latency, cost-efficient chat

$0.30/1M in · $1.20/1M out ↓46% View →
Veo 3.1 Video

Veo 3.1

veo-3.1

Text / image to video · multiple tiers · multiple resolutions

$0.075/clip+ View →
Seedance 2.5 Video

Seedance 2.5

seedance-2.5

Text / image to video · up to 30 seconds

$0.158/s View →
Seedance 2.0 Video

Seedance 2.0

seedance-2.0

Text / image to video · 5-15 seconds

$0.150/s View →
Seedance 2.0 Fast Video

Seedance 2.0 Fast

seedance-2.0-fast

Low latency · lowest cost

$0.100/s View →
Seedance 2.0 Mini Video

Seedance 2.0 Mini

seedance-2.0-mini

Entry tier · lowest cost

$0.066/s View →
Wan 3.0 Video

Wan 3.0

wan3.0-video

One endpoint, five modes · up to 30 seconds

$0.120/s View →
Wan 3.0 Prime Video

Wan 3.0 Prime

wan3.0-video-prime

Same capabilities · several times faster

$0.160/s View →
MiniMax H3 Video

MiniMax H3

minimax-h3

1080P · native audio

$0.036/s View →
Grok Imagine Video 1.5 Video

Grok Imagine Video 1.5

grok-imagine-video-1.5

Billed per clip · up to 15 seconds

$0.300/clip+ View →