GPT-6 Luna
⧉GPT-6 Luna is the most efficient model in the GPT-6 family, built for focused, high-volume tasks: the same 1,050,000-token context and up to 128,000 output tokens, reasoning_effort in six steps from none to max, image input and tool calling. Fully OpenAI-SDK compatible - set model to gpt-6-luna - and served on both /chat/completions and /responses. Pay-as-you-go, failed calls never billed; cached input bills at one tenth.
Live Test · Playground
Try out GPT-6 Luna right here (available after login).
Input
Advanced
This model is served over a ChatGPT-subscription upstream that does not accept temperature / top_p / max_tokens (they are silently ignored). Use reasoning_effort below for thinking depth, and prompt wording for length.
Conversation
About GPT-6 Luna
GPT-6 Luna is the most efficient model in OpenAI's GPT-6 family, built for focused, high-volume tasks and served on NezhaGate through the OpenAI-compatible API: the same 1,050,000-token context, up to 128,000 output tokens per call, reasoning_effort in six steps from none to max, image input and tool calling, on both /chat/completions and /responses. Set model to gpt-6-luna and you are connected. Pay-as-you-go, failed calls are never billed, and cached input settles at one tenth of the input rate.
Use cases
Support replies, intent detection and tag extraction at scale, where cost and latency come first.
The many small judgement and formatting steps in an agent flow: Luna saves both time and budget.
A 1,050,000-token context at a very low unit price suits bulk summarising, extraction and rewriting of long documents.
Get the whole flow running on Luna first, then promote the critical steps to GPT-6 Sol or Astra.
How to choose
Send simple, high-volume, latency-sensitive work to GPT-6 Luna; switch to GPT-6 Sol for stronger coding and reasoning and to GPT-6 Astra for the hardest tasks. All three belong to the GPT-6 family, so switching is a one-field change.
FAQ
How is it different from GPT-6 Sol?
Which reasoning_effort levels are supported?
Does it support streaming?
Am I charged for failed calls?
Related Models
Explore other models you can integrate.
GPT-6 Sol
Built for complex coding and agentic workflows - 1.05M context
GPT-6 Astra
Next-generation OpenAI flagship - 1.05M context
GPT-5.6 Luna
Fast, economical agentic coding model
GPT-5.6 Terra
Balanced everyday agentic coding model
Gemini 3.8 Flash
Latest Flash · adaptive thinking