# Claude Sonnet 5.5 — /chat/completions

OpenAI-compatible chat completions backed by Anthropic Claude Sonnet 5.5, the newest Sonnet and the best combination of speed and intelligence: a 1M-token context window, up to 128K output tokens per request and adaptive thinking. Image input, tool calling, prompt caching, stop sequences and streaming are supported; set model to claude-sonnet-5-5. On the OpenAI-compatible endpoint the model thinks at its default effort and returns no thinking content, and response_format is not guaranteed. For effort control (output_config.effort: low / medium / high / xhigh / max), thinking summaries (thinking.display: summarized), structured output (output_config.format) or PDF input, use the native Anthropic Messages API at /anthropic/v1/messages, which Claude Code connects to directly. temperature / top_p / top_k are ignored, a forced tool choice is treated as auto, and thinking set to disabled is ignored.

**Endpoint:** `POST https://nezhagate.com/v1/chat/completions`

## Authentication
```
Authorization: Bearer YOUR_API_KEY
Content-Type: application/json
```

## Request body
| Parameter | Type | Required | Description |
| --- | --- | --- | --- |
| `model` | string | Yes | Model ID, here claude-sonnet-5-5. |
| `messages` | array | Yes | Array of messages; each has role (system/user/assistant) and content. content may be a string, or an array of {type:text} and {type:image_url} parts for image understanding (multimodal/vision). |
| `stream` | boolean | No | Stream the response as SSE. Default false. |
| `temperature` | number | No | No effect: this model ignores temperature / top_p (no error) and samples at its default. |
| `max_tokens` | integer | No | Maximum number of tokens to generate. |
| `web_search` | boolean | No | Set true to enable web search: the gateway augments the prompt with live results (citing sources) before the model answers. Can also be triggered via a tools entry {"type":"web_search"}. |
| `reasoning_effort` | string | No | No effect on the OpenAI-compatible endpoint: the model thinks at its default effort. To set the effort, call the native Anthropic API and send output_config.effort (low / medium / high / xhigh / max). |

## Request example
```bash
curl https://nezhagate.com/v1/chat/completions -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d '{"model": "claude-sonnet-5-5", "messages": [{"role": "user", "content": "Hello"}], "stream": false}'
```

## Response
```json
{
  "id": "chatcmpl_xxx",
  "object": "chat.completion",
  "model": "claude-sonnet-5-5",
  "choices": [
    {"index": 0, "message": {"role": "assistant", "content": "Hello!"}, "finish_reason": "stop"}
  ],
  "usage": {"prompt_tokens": 11, "completion_tokens": 7, "total_tokens": 18}
}
```

## Image input (Vision)
Put an image in the message content array and the model will analyze it (visual Q&A, reading text / OCR, …). image_url accepts a public image link or an inline base64 data URL (data:image/png;base64,...). Available on multimodal models (gpt-5.5, gemini series, …).

```bash
curl https://nezhagate.com/v1/chat/completions -H "Authorization: Bearer YOUR_API_KEY" -H "Content-Type: application/json" -d '{"model": "claude-sonnet-5-5", "messages": [{"role": "user", "content": [{"type": "text", "text": "What is in this image?"}, {"type": "image_url", "image_url": {"url": "https://example.com/photo.jpg"}}]}]}'
```