NezhaGateNezhaGate
Chat

GPT-6 Luna

⧉
gpt-6-luna
Model Type AstraSolLuna

GPT-6 Luna is the most efficient model in the GPT-6 family, built for focused, high-volume tasks: the same 1,050,000-token context and up to 128,000 output tokens, reasoning_effort in six steps from none to max, image input and tool calling. Fully OpenAI-SDK compatible - set model to gpt-6-luna - and served on both /chat/completions and /responses. Pay-as-you-go, failed calls never billed; cached input bills at one tenth.

ChatGPT-6 familyFastLow cost1.05M context

Live Test · Playground

Try out GPT-6 Luna right here (available after login).

Input

Advanced

This model is served over a ChatGPT-subscription upstream that does not accept temperature / top_p / max_tokens (they are silently ignored). Use reasoning_effort below for thinking depth, and prompt wording for length.

Web search
Memory · multi-turn

Conversation

Start a conversationType a message below to begin

About GPT-6 Luna

GPT-6 Luna is the most efficient model in OpenAI's GPT-6 family, built for focused, high-volume tasks and served on NezhaGate through the OpenAI-compatible API: the same 1,050,000-token context, up to 128,000 output tokens per call, reasoning_effort in six steps from none to max, image input and tool calling, on both /chat/completions and /responses. Set model to gpt-6-luna and you are connected. Pay-as-you-go, failed calls are never billed, and cached input settles at one tenth of the input rate.

Use cases

High-volume chat and classification

Support replies, intent detection and tag extraction at scale, where cost and latency come first.

Simple steps inside agent flows

The many small judgement and formatting steps in an agent flow: Luna saves both time and budget.

Bulk long-document work

A 1,050,000-token context at a very low unit price suits bulk summarising, extraction and rewriting of long documents.

Fast prototyping

Get the whole flow running on Luna first, then promote the critical steps to GPT-6 Sol or Astra.

How to choose

Send simple, high-volume, latency-sensitive work to GPT-6 Luna; switch to GPT-6 Sol for stronger coding and reasoning and to GPT-6 Astra for the hardest tasks. All three belong to the GPT-6 family, so switching is a one-field change.

FAQ

How is it different from GPT-6 Sol?
Luna is faster and cheaper, built for high-volume, low-latency traffic; Sol is stronger at complex coding and agentic workflows. Context and integration are identical.
Which reasoning_effort levels are supported?
Six: none / low / medium / high / xhigh / max. When omitted the gateway requests medium, the model's own default; more thinking means more output tokens, billed at actual usage.
Does it support streaming?
Yes. Pass stream=true and the response arrives as SSE chunks.
Am I charged for failed calls?
No. Only calls that return a proper response are settled at actual usage; upstream errors and timeouts never touch your balance.

Related Models

Explore other models you can integrate.

View all →