GPT-6 Astra
⧉GPT-6 Astra is the next-generation flagship OpenAI released in September 2026: a 1,050,000-token context window, up to 128,000 output tokens, reasoning_effort up to xhigh, image input, tool calling and prompt caching. Fully OpenAI-SDK compatible - set model to gpt-6-astra - and served on both /chat/completions and /responses. Pay-as-you-go, failed calls never billed; cached input bills at one tenth, and we do not add the long-context surcharge the official API applies.
Live Test · Playground
Try out GPT-6 Astra right here (available after login)。
Input
Advanced
This model is served over a ChatGPT-subscription upstream that does not accept temperature / top_p / max_tokens (they are silently ignored). Use reasoning_effort below for thinking depth, and prompt wording for length.
Conversation
About GPT-6 Astra
GPT-6 Astra is the next-generation flagship OpenAI released in September 2026, served on NezhaGate through the OpenAI-compatible API: a 1,050,000-token context window, up to 128,000 output tokens, reasoning_effort for thinking depth (up to xhigh), image input, tool calling and prompt caching, on both /chat/completions and /responses. Set model to gpt-6-astra and keep your SDK and request shape. Pay-as-you-go, failed calls never billed, and cached input settles at one tenth of the input rate.
Use cases
1,050,000 tokens holds a whole repository, a contract set or a long report in one call, with no chunking and stitching of your own.
It speaks the OpenAI-compatible endpoint, so Codex, Cursor and Claude Code switch to GPT-6 Astra by changing the model field alone.
reasoning_effort goes up to xhigh, so you can spend the thinking budget on the hardest part and keep simple requests on low.
Image input with text output: screenshot understanding, chart questions and document parsing all use the same endpoint.
How to choose
Pick GPT-6 Astra when you need the strongest reasoning or a very long context; gpt-5.6-terra is cheaper for everyday development and gpt-5.6-luna for high-volume low-latency work. All three use the OpenAI-compatible endpoint, so switching is a one-field change.
FAQ
What are the context and output limits?
How do I get cache hits?
How does the price compare with the official API?
Is reasoning_effort supported?
Can I use /v1/responses?
Related Models
Explore other models you can integrate.
GPT-5.6 Sol
Frontier flagship for agentic coding
GPT-5.6 Terra
Balanced everyday agentic coding model
GPT-5.6 Luna
Fast, economical agentic coding model
GPT-5.5
Flagship chat and reasoning model
GPT Image 2
High-quality text-to-image / image-to-image model
GPT Image 2.5 Flare
Next-gen image model - clean and smooth
GPT Image 2.5 Sunburst
Next-gen image model - richer texture
Nano Banana 2
High-quality text-to-image / image-to-image model
Nano Banana Pro
Flagship text-to-image / image-to-image model
Claude Sonnet 4.6
Balanced, efficient chat and coding model
Claude Opus 5
Next-generation flagship from Anthropic
Claude Fable 5
Anthropic Fable series · narrative and long-form writing
Gemini 3.1 Pro
Next-generation flagship reasoning model from Google
Gemini 3.8 Flash
Latest Flash · adaptive thinking
Gemini 3.7 Flash
Previous Flash · adaptive thinking
Gemini 3.6 Flash
Next-generation fast thinking model · four thinking budgets
Gemini 3.6 Flash High
Deep thinking tier
Gemini 3.6 Flash Low
Fastest, lowest-cost tier
Gemini 3.6 Flash Tiered
Auto-tiered thinking
Gemini 3 Flash
Fast, low-latency, cost-efficient chat model
Gemini 3.5 Flash
Next-generation fast thinking model
Gemini 2.5 Flash
Fast, low-latency, cost-efficient chat
Veo 3.1
Text / image to video · multiple tiers · multiple resolutions
Seedance 2.5
Text / image to video · up to 30 seconds
Seedance 2.0
Text / image to video · 5-15 seconds
Seedance 2.0 Fast
Low latency · lowest cost
Seedance 2.0 Mini
Entry tier · lowest cost
Wan 3.0
One endpoint, five modes · up to 30 seconds
Wan 3.0 Prime
Same capabilities · several times faster
MiniMax H3
1080P · native audio
Grok Imagine Video 1.5
Billed per clip · up to 15 seconds