AI Model Market
Browse every available AI model, compare pricing and capabilities, and integrate in minutes. Currently offering 21 models, with more added continuously. Plus 1 not yet open.

GPT-5.6
The GPT-5.6 family: Sol is the frontier flagship, Terra the balanced everyday model, Luna the fast and economical one -- a single family spanning frontier intelligence to high-speed execution. Use Model Type to pick a tier.

GPT-5.5
GPT-5.5 delivers industry-leading chat, reasoning and coding, with streaming, long context and structured tasks. Fully OpenAI-SDK compatible: plug it into Codex, Cursor, Claude Code and similar tools with zero code changes.

GPT-6 Astra
GPT-6 Astra is the next-generation flagship OpenAI released in September 2026: a 1,050,000-token context window, up to 128,000 output tokens, reasoning_effort up to xhigh, image input, tool calling and prompt caching. Fully OpenAI-SDK compatible - set model to gpt-6-astra - and served on both /chat/completions and /responses. Pay-as-you-go, failed calls never billed; cached input bills at one tenth, and we do not add the long-context surcharge the official API applies.

Claude Sonnet 4.6
Claude Sonnet 4.6 strikes an excellent balance between capability and speed, excelling at everyday chat, code and agents at a great price. Fully OpenAI-API compatible: set model to claude-sonnet-4-6 and you are connected.

Claude Opus 5
Claude Opus 5 is the next-generation flagship from Anthropic, with adaptive thinking on by default. It excels at complex reasoning, code engineering and agent orchestration, and supports image understanding, tool use and prompt caching. Fully OpenAI-API compatible: set model to claude-opus-5; or use the native Anthropic Messages API so Claude Code connects directly.

Claude Fable 5
Claude Fable 5 belongs to the Anthropic Fable series: same generation and capability base as Opus 5 (adaptive thinking, tool use, image understanding, prompt caching), with a finer touch and richer emotional layering in Chinese narrative and long-form writing. Fully OpenAI-API compatible: set model to claude-fable-5; the native Anthropic Messages API is supported as well.

Gemini 3.1 Pro
Gemini 3.1 Pro is the next-generation flagship multimodal model from Google, with a very long context window and top-tier reasoning and coding. Fully OpenAI-API compatible: set model to gemini-3.1-pro in your existing SDK. Low-reasoning tier by default (faster and cheaper); pass reasoning_effort=high for the high tier (billed at the Preview rate). Responses carry a reasoning_content field so you can see the thinking (streamed chunk by chunk as well).

Gemini 3.8 Flash
Gemini 3.8 Flash is the latest Flash generation from Google (released September 2026), built for long-running coding, agent workflows and complex reasoning: 1M-token context, up to 64K output, image input. It uses a dynamic thinking budget -- quick answers for simple questions, more thinking for hard ones. Fully OpenAI-API compatible: set model to gemini-3.8-flash.

Gemini 3.7 Flash
Gemini 3.7 Flash is the generation before 3.8, with the same dynamic thinking budget, 1M-token context and image input. A good fit if your prompts are already tuned for 3.7 and you do not want to switch yet. Fully OpenAI-API compatible: set model to gemini-3.7-flash.

Gemini 3.6 Flash
The Gemini 3.6 Flash family: one model, four thinking budgets -- Medium for everyday balance, High for deep reasoning, Low for speed and low cost, Tiered lets the model pick by difficulty. All four cost the same; more thinking simply means more output tokens. 1M context, image input. Use Model Type to pick a tier.

Gemini 3 Flash
Gemini 3 Flash focuses on speed and cost efficiency, for high-volume, low-latency chat and agent workloads. Fully OpenAI-API compatible with streaming: set model to gemini-3-flash-preview.

Gemini 3.5 Flash
Gemini 3.5 Flash is the next-generation fast model from Google with built-in thinking: fast and cost-efficient, for high-volume chat and agent workloads. Fully OpenAI-API compatible: set model to gemini-3.5-flash.

Gemini 2.5 Flash
Gemini 2.5 Flash focuses on speed and cost efficiency, for large-scale, low-latency chat and agent workloads, with multimodal input. Fully OpenAI-API compatible: set model to gemini-2.5-flash.








