Documentation
LMHaven is a single REST surface for chat, vision, image, voice, and video generation — flat 5× cheaper than first-party list across every model. Two API shapes are supported side-by-side: OpenAI Chat Completions and Anthropic Messages. Drop in by changing one base URL.
Quickstart
- Create an API key at dashboard/keys — it starts with
sk-lmh-. - Point your favourite SDK at
https://lmhaven.xyz/api/v1as the base URL. - Make a request. Same key works across chat, image, audio, and video.
curl https://lmhaven.xyz/api/v1/chat/completions \
-H "Authorization: Bearer sk-lmh-…" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-7",
"messages": [{"role":"user","content":"Write a haiku about debugging."}]
}'Authentication
Every endpoint accepts the API key on either header — pick whichever your SDK already sends:
Authorization: Bearer sk-lmh-…— used by OpenAI SDKs, Cursor, Continue, Cline, Roo Code, and any tool pointed at/api/v1.x-api-key: sk-lmh-…— used by Anthropic SDKs and Claude Code CLI when pointed at/v1/messages.
Keys are workspace-scoped. Revoke from dashboard/keys — subsequent calls fail with 401 immediately, no propagation delay.
Tool integrations
Drop-in configs for the most-asked tools.
Cursor
Settings → Models → OpenAI API Key:
- API Key:
sk-lmh-… - Override OpenAI Base URL:
https://lmhaven.xyz/api/v1 - Add custom model:
gemini-3-1-pro,glm-5-1, etc. - Claude & GPT in Cursor: Cursor special-cases model ids beginning with
claude-,gpt-, oro<n>and routes them through its own provider logic (ignoring this base URL), so they fail. Use thelmh--prefixed aliases instead, which Cursor treats as plain custom models:lmh-claude-opus-4-7,lmh-claude-sonnet-4-6,lmh-gpt-5-5,lmh-gpt-5-2,lmh-o4-mini. Same models, same pricing.
Continue.dev
Continue uses YAML (~/.continue/config.yaml) now — the older config.json format is deprecated. Use provider: openai for every model (chat models route through the OpenAI-compatible endpoint, including Claude).
name: LMHaven
version: 0.0.1
schema: v1
models:
- name: Claude Opus 4.7
provider: openai
model: claude-opus-4-7
apiBase: https://lmhaven.xyz/api/v1
apiKey: sk-lmh-…
roles:
- chat
- edit
- apply
- name: Kimi K2.6
provider: openai
model: kimi-k2-6
apiBase: https://lmhaven.xyz/api/v1
apiKey: sk-lmh-…
roles:
- chat
- editCline
Cline → Settings → API Provider: OpenAI Compatible
- Base URL:
https://lmhaven.xyz/api/v1 - API Key:
sk-lmh-… - Model: any from the catalog (
claude-sonnet-4-6recommended).
Roo Code
Same as Cline. Provider type: OpenAI Compatible.
Aider
Aider supports any OpenAI-compatible endpoint. Two env vars + the model flag:
export OPENAI_API_BASE="https://lmhaven.xyz/api/v1" export OPENAI_API_KEY="sk-lmh-…" aider --model openai/claude-sonnet-4-6 # or any chat model id from /v1/models
Droid (Factory.ai)
In Droid's BYOK config, add each model with:
- Provider type:
generic-chat-completion-api(oropenai) - Base URL:
https://lmhaven.xyz/api - API key: your
sk-lmh-… - Model name: any chat model id from /v1/models
Tool use works on every chat model — Droid will see tool_calls the same way it does on OpenAI directly. No need to set provider: "anthropic" for Claude / Gemini models; one config covers every chat model.
Claude Code CLI
Claude Code speaks the Anthropic Messages API. Claude Code appends /v1/messages to the base URL, so point the base at /api (it resolves to /api/v1/messages). Two env vars:
export ANTHROPIC_BASE_URL="https://lmhaven.xyz/api" export ANTHROPIC_AUTH_TOKEN="sk-lmh-…" # Some Claude Code builds also read ANTHROPIC_API_KEY: export ANTHROPIC_API_KEY="sk-lmh-…" claude # any claude-* model id resolves automatically (opus / sonnet / haiku tier)
API reference
GET /v1/models
OpenAI-shaped catalog discovery. Public — no auth needed.
curl https://lmhaven.xyz/api/v1/models
{
"object": "list",
"data": [
{ "id": "claude-opus-4-7", "object": "model", "owned_by": "lmhaven", … },
{ "id": "gemini-3-1-pro", … }
]
}POST /v1/chat/completions
OpenAI-compatible. Same shape as platform.openai.com/docs/api-reference/chat/create.
Supported fields: model, messages, max_tokens, system, stream, temperature, top_p, stop (string or array, capped at 5), tools, tool_choice, reasoning_effort, thinking_budget.
Reasoning control for thinking-capable models (Gemini 2.5+/3.x, GLM, etc.). Set reasoning_effort to one of "none", "low", "medium", "high" for preset budgets, or pass an explicit integer via thinking_budget when you need fine-grained control (Antigravity-style). For Gemini the preset map is none=0, low=128, medium=2048, high=16384; Gemini Pro clamps to its 128 minimum when thinking is enabled. OpenAI's "minimal" alias is accepted and treated as "none".
Multimodal content arrays ([{type:"text",…}, {type:"image_url",…}, {type:"input_audio",…}]) are fully supported. Pass image_url.url as either an https://… URL (we fetch and inline server-side, SSRF-hardened) or a data:image/…;base64,… URI (parsed locally). PDFs and video go through the same field — infer-by-extension picks the right size cap and MIME allowlist.
Tool use: tools, tool_choice, role:"tool" messages, and assistant.tool_calls all work across every chat model — Claude, GPT, Gemini, GLM, Grok, DeepSeek, Qwen, Kimi, Llama, Mistral. Streaming emits OpenAI-shape tool_calls deltas. System prompts work either via top-level system or as a role:"system" message — both paths are normalized.
Rejected with 400 (loud, not silent): response_format (json_mode / structured outputs), n > 1.
curl https://lmhaven.xyz/api/v1/chat/completions \
-H "Authorization: Bearer sk-lmh-…" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"messages": [
{"role": "system", "content": "You are a precise SQL tutor."},
{"role": "user", "content": "Explain window functions in two sentences."}
],
"stream": true,
"temperature": 0.4
}'POST /v1/messages
Anthropic-compatible. Same shape as docs.anthropic.com/en/api/messages. Use this for vision, tool use, and any Claude SDK consumer (including Claude Code CLI).
Full feature parity: messages (string or content blocks), system (string or text-block array), max_tokens, temperature, stream, tools, tool_choice, thinking. Image content blocks ({type:"image",…}) and tool_use / tool_result blocks all flow through.
Extended thinking in Anthropic shape — pass "thinking": { "type": "enabled", "budget_tokens": 4096 } to enable, or "thinking": { "type": "disabled" } / null to turn it off. When the request routes to Gemini Pro this is mapped to thinkingConfig.thinkingBudget with a floor of 128. For Claude SKUs the field is forwarded verbatim to the upstream Anthropic-shape API.
Model aliases: versioned Anthropic ids are accepted by tier — claude-sonnet-4-5, claude-3-5-sonnet-20241022, claude-haiku-4-5-latest all resolve to the matching tier in our catalog (claude-opus-4-7, claude-sonnet-4-6, claude-haiku-4-5), so SDKs that pin to a specific Anthropic version work out of the box.
curl https://lmhaven.xyz/api/v1/messages \
-H "x-api-key: sk-lmh-…" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-7",
"max_tokens": 1024,
"messages": [{"role":"user","content":"hi"}]
}'POST /v1/images/generations
OpenAI-shaped. Returns 24-hour presigned URLs.
Supported: model, prompt, size, aspect_ratio, images (1–10 reference images for editing / composition), n (1–4 outputs per call; each billed at the full per-image rate).
Rejected with 400: quality, style, response_format != "url".
curl https://lmhaven.xyz/api/v1/images/generations \
-H "Authorization: Bearer sk-lmh-…" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2",
"prompt": "A watercolor city skyline at dusk, cinematic lighting",
"aspect_ratio": "16:9",
"n": 4
}'With reference images. Pass images as an array of data URLs (or raw base64) — they become inline-image parts on the generation request for editing / composition. Up to 10 reference images per call.
curl https://lmhaven.xyz/api/v1/images/generations \
-H "Authorization: Bearer sk-lmh-…" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2",
"prompt": "Place this character on a snowy mountain at golden hour.",
"aspect_ratio": "16:9",
"images": [
"data:image/png;base64,iVBORw0KGgoAAAANSUhEUg…",
"data:image/jpeg;base64,/9j/4AAQSkZJRgABAQEAYABg…"
],
"n": 2
}'POST /v1/audio/speech
OpenAI-shaped TTS. Returns audio bytes inline.
Voice aliases (OpenAI names): alloy · echo · fable · onyx · nova · shimmer · coral · ash · ballad · sage · verse — each maps to the closest stock ElevenLabs voice we host.
ElevenLabs voice names: rachel · adam · bella · antoni · domi · elli · josh · arnold · sam. Or pass any raw ElevenLabs voice id directly.
Formats: mp3 (default), pcm, wav. opus/aac/flac rejected with 400.
curl https://lmhaven.xyz/api/v1/audio/speech \
-H "Authorization: Bearer sk-lmh-…" \
-H "Content-Type: application/json" \
-o output.mp3 \
-d '{
"model": "elevenlabs-v3-multilingual",
"voice": "alloy",
"input": "Hello from lmhaven."
}'POST /v1/video/generations
Submit prompt → block until ready → 24-hour presigned URLs. Veo 3 / Veo 3 Fast emit native audio in the clip. Generation takes 30–90s typical, 5-minute hard ceiling.
Supported: model, prompt, ratio, duration, resolution, image (first-frame reference, Veo only), end_image (last-frame reference for frames-to-video transition), n (1–4 videos per call; each billed at the full per-video rate).
curl https://lmhaven.xyz/api/v1/video/generations \
-H "Authorization: Bearer sk-lmh-…" \
-H "Content-Type: application/json" \
-d '{
"model": "veo-3.1",
"prompt": "A calm ocean wave at sunset, slow camera pan",
"ratio": "16:9",
"duration": 8,
"resolution": "720p",
"n": 1
}'Image-to-video. Pass image as a data URL — Veo animates from that still as the first frame. Add end_image to lock the last frame too (Veo generates a transition between the two stills).
curl https://lmhaven.xyz/api/v1/video/generations \
-H "Authorization: Bearer sk-lmh-…" \
-H "Content-Type: application/json" \
-d '{
"model": "veo-3.1",
"prompt": "Subject walks toward the camera, slow zoom in",
"ratio": "16:9",
"duration": 8,
"image": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUg…",
"end_image": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUg…"
}'{
"data": [
{ "url": "https://cdn.lmhaven.xyz/videos/…/abc.mp4?…", "duration_seconds": 8 },
{ "url": "https://cdn.lmhaven.xyz/videos/…/def.mp4?…", "duration_seconds": 8 }
]
}POST /v1/files
Upload media once, reference it from many chat calls by file_id. The OpenAI and Anthropic SDKs both call this path automatically when you pass a File handle to a chat method — drop-in for either ecosystem. Multipart file + purpose body.
Limits: 200 files per account, 2 GB total, 30-day TTL (lazy-expired on read past TTL). Per-file caps mirror inline upload — 20 MB image / video / audio, 32 MB PDF. Same MIME allowlist applies.
curl https://lmhaven.xyz/api/v1/files \ -H "Authorization: Bearer sk-lmh-…" \ -F purpose=vision \ -F file=@/path/to/photo.png
Models & pricing
Flat 5× cheaper than first-party across every model — no tiered SKUs, no cache discount math by default. Rates verified weekly against each provider’s published list price (sources cited inline in lib/models.ts).
Text · vision (per 1M tokens, input / output)
| Model | LMHaven | First-party |
|---|---|---|
| claude-fable-5 | $2.00/1M in · $10.00/1M out | Anthropic $10.00/1M / $50.00/1M |
| claude-opus-4-7 | $1.00/1M in · $5.00/1M out | Anthropic $5.00/1M / $25.00/1M |
| claude-opus-4-8 | $1.00/1M in · $5.00/1M out | Anthropic $5.00/1M / $25.00/1M |
| claude-opus-5 | $1.00/1M in · $5.00/1M out | Anthropic $5.00/1M / $25.00/1M |
| claude-sonnet-5 | $0.60/1M in · $3.00/1M out | Anthropic $3.00/1M / $15.00/1M |
| claude-sonnet-4-6 | $0.60/1M in · $3.00/1M out | Anthropic $3.00/1M / $15.00/1M |
| claude-haiku-4-5 | $0.20/1M in · $1.00/1M out | Anthropic $1.00/1M / $5.00/1M |
| gpt-5-2 | $1.00/1M in · $8.00/1M out | OpenAI $5.00/1M / $40.00/1M |
| gpt-5-4 | $1.00/1M in · $8.00/1M out | OpenAI $5.00/1M / $40.00/1M |
| gpt-5-5 | $1.25/1M in · $10.00/1M out | OpenAI $6.25/1M / $50.00/1M |
| gpt-5-6-sol | $1.00/1M in · $6.00/1M out | OpenAI $5.00/1M / $30.00/1M |
| gpt-5-6-terra | $0.50/1M in · $3.00/1M out | OpenAI $2.50/1M / $15.00/1M |
| gpt-5-6-luna | $0.20/1M in · $1.20/1M out | OpenAI $1.00/1M / $6.00/1M |
| gpt-5-mini | $0.05/1M in · $0.40/1M out | OpenAI $0.25/1M / $2.00/1M |
| o4-mini | $0.22/1M in · $0.88/1M out | OpenAI $1.10/1M / $4.40/1M |
| gemini-3-1-pro | $0.50/1M in · $4.00/1M out | Google $2.50/1M / $20.00/1M |
| gemini-3-5-flash | $0.06/1M in · $0.50/1M out | Google $0.30/1M / $2.50/1M |
| gemini-3-1-flash-lite | $0.02/1M in · $0.08/1M out | Google $0.10/1M / $0.40/1M |
| llama-4-maverick | $0.07/1M in · $0.28/1M out | Meta $0.35/1M / $1.40/1M |
| llama-4-scout | $0.04/1M in · $0.16/1M out | Meta $0.20/1M / $0.80/1M |
| deepseek-r1 | $0.11/1M in · $0.44/1M out | DeepSeek $0.55/1M / $2.19/1M |
| deepseek-v3-2 | $0.05/1M in · $0.22/1M out | DeepSeek $0.27/1M / $1.10/1M |
| qwen-3-coder | $0.20/1M in · $0.90/1M out | Alibaba $1.00/1M / $4.50/1M |
| qwen-3-next | $0.03/1M in · $0.12/1M out | Alibaba $0.15/1M / $0.60/1M |
| kimi-k2-thinking | $0.12/1M in · $0.50/1M out | Moonshot $0.60/1M / $2.50/1M |
| kimi-k2-6 | $0.15/1M in · $0.60/1M out | Moonshot $0.75/1M / $3.00/1M |
| kimi-k2-7-code | $0.19/1M in · $0.80/1M out | Moonshot $0.95/1M / $4.00/1M |
| deepseek-v4-pro | $0.10/1M in · $0.50/1M out | DeepSeek $0.50/1M / $2.50/1M |
| deepseek-v4-flash | $0.04/1M in · $0.16/1M out | DeepSeek $0.20/1M / $0.80/1M |
| grok-4-3 | $0.80/1M in · $4.00/1M out | xAI $4.00/1M / $20.00/1M |
| gpt-oss-120b | $0.05/1M in · $0.20/1M out | OpenAI $0.25/1M / $1.00/1M |
| grok-4-1-fast | $0.20/1M in · $1.00/1M out | xAI $1.00/1M / $5.00/1M |
| grok-4-2 | $0.60/1M in · $3.00/1M out | xAI $3.00/1M / $15.00/1M |
| codestral-2 | $0.06/1M in · $0.18/1M out | Mistral $0.30/1M / $0.90/1M |
| mistral-medium-3 | $0.08/1M in · $0.40/1M out | Mistral $0.40/1M / $2.00/1M |
Image (per call)
| Model | LMHaven | First-party |
|---|---|---|
| nano-banana-pro | $0.008 | Google $0.039 |
| nano-banana-2 | $0.027 | Google $0.134 |
Voice (per 1K characters)
| Model | LMHaven | First-party |
|---|
Video (per 8s clip)
| Model | LMHaven | First-party |
|---|---|---|
| veo-3.1 | $0.640 | Google $3.200 |
| veo-3.1-fast | $0.160 | Google $0.800 |
| veo-3 | $0.640 | Google $3.200 |
| veo-3-fast | $0.160 | Google $0.800 |
| veo-2 | $0.800 | Google $4.000 |
Reseller / enterprise rates: per-account discount factors are available (8×, 10×, 20× cheaper than list). Contact us via Discord for contracted tiers.
Multimodal input
Both endpoints accept image, video, audio, and PDF input on every vision-capable model (claude-*, gemini-3-*). Two ref forms, both fully supported:
- Inline base64 — Anthropic
{type:"image", source:{type:"base64",media_type,data}}or OpenAIdata:image/png;base64,…URIs. Parsed locally. - URL (https only) — Anthropic
{type:"image", source:{type:"url",url}}or OpenAIimage_url.url: "https://…". We fetch the URL server-side with an SSRF guard (RFC1918 / link-local / cloud-metadata addresses are pre-rejected by DNS resolution; redirects re-validated per hop, capped at 3) and inline as base64 before forwarding.
Limits: 20 MB image / video / audio, 32 MB PDF, 10 s fetch timeout. Allowed MIMEs: PNG / JPEG / GIF / WEBP / HEIC / HEIF for image; MP4 / WebM / MPEG / MOV / AVI / 3GPP for video; MP3 / WAV / AAC / M4A / OGG / FLAC for audio; PDF for documents. Anthropic document blocks (PDF + plain-text + structured-content forms) are accepted on /v1/messages.
Files API references are also live — upload once via POST /v1/files, then pass the resulting file_id as {type:"file", file:{file_id}} on /v1/chat/completions or as {type:"image", source:{type:"file", file_id}} on /v1/messages. Per-account isolation is enforced server-side — a file_id from another account returns 404.
Anthropic shape — image (URL)
curl https://lmhaven.xyz/api/v1/messages \
-H "x-api-key: sk-lmh-…" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"max_tokens": 512,
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{
"type": "image",
"source": {
"type": "url",
"url": "https://example.com/photo.png"
}
}
]
}]
}'Anthropic shape — image (base64)
# base64-encode locally first
B64=$(base64 -i photo.png)
curl https://lmhaven.xyz/api/v1/messages \
-H "x-api-key: sk-lmh-…" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d "{
\"model\": \"claude-sonnet-4-6\",
\"max_tokens\": 512,
\"messages\": [{
\"role\": \"user\",
\"content\": [
{\"type\": \"text\", \"text\": \"What is in this image?\"},
{
\"type\": \"image\",
\"source\": {
\"type\": \"base64\",
\"media_type\": \"image/png\",
\"data\": \"$B64\"
}
}
]
}]
}"Anthropic shape — PDF document
curl https://lmhaven.xyz/api/v1/messages \
-H "x-api-key: sk-lmh-…" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4-7",
"max_tokens": 1024,
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Summarize this contract in three bullets."},
{
"type": "document",
"source": {"type": "url", "url": "https://example.com/contract.pdf"},
"title": "Master Services Agreement",
"context": "Reviewing for renewal terms"
}
]
}]
}'OpenAI shape — image_url + input_audio
curl https://lmhaven.xyz/api/v1/chat/completions \
-H "Authorization: Bearer sk-lmh-…" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-4-6",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image."},
{"type": "image_url", "image_url": {"url": "https://example.com/photo.png"}}
]
}]
}'Files API — upload once, reuse by file_id
# Returns {"id": "file_…", "object": "file", "bytes": …, "filename": …, "mime_type": …}
curl https://lmhaven.xyz/api/v1/files \
-H "Authorization: Bearer sk-lmh-…" \
-F purpose=vision \
-F file=@photo.pngTool use (function calling)
Works on both endpoints, across every chat model in the catalog — Claude, GPT, Gemini, GLM, Grok, DeepSeek, Qwen, Kimi, Llama. No per-model configuration needed. Whichever shape your client speaks, point it at LMHaven and tools just work.
- OpenAI shape on /v1/chat/completions —
tools,tool_choice,role:"tool",assistant.tool_calls. Streaming emitstool_callsdeltas. Used by Cursor, Continue, Cline, Roo Code, Aider, Droid, OpenAI SDKs. - Anthropic shape on /v1/messages —
toolsarray,tool_useresponse blocks,tool_resultin user follow-ups,input_json_deltastreaming. Used by Claude Code CLI, Anthropic SDKs.
Translation between OpenAI and Anthropic tool shapes happens server-side and is transparent — your client sees the shape it asked for, no matter which model serves the request.
Streaming
OpenAI shape (/v1/chat/completions)
SSE with data: {chat.completion.chunk} frames, terminated by data: [DONE]. Each chunk carries choices[0].delta.content for text, or choices[0].delta.tool_calls for function-call deltas. Final chunk has choices[0].finish_reason set to "stop" for plain text or "tool_calls" when the model is asking to call a function, plus a usage object.
Anthropic shape (/v1/messages)
Full Anthropic event sequence: message_start → content_block_start → content_block_delta×N → content_block_stop → message_delta → message_stop. Tool-use blocks emit input_json_delta with the full args.
Prompt caching
Pass cache_control headers in your request — upstream honours them for latency wins regardless of any LMHaven setting.
Cache discounts on the bill are an opt-in toggle on /dashboard/keys. With it enabled:
- Anthropic cache reads: 10% of input rate (~50× off list net).
- Anthropic cache writes: 125% of input rate (one-time premium per cached prefix; break-even after ~2 reads).
- Gemini cache reads: 25% of input rate (~20× off list net).
- Gemini cache writes: 100% of input rate (no premium).
- Output tokens: unaffected, always at your standard rate.
Quality is identical with the toggle on or off. Caching only changes billing — the model still runs full inference on every call. Caching reuse is at the prefix-processing layer, not response replay.
Errors
| Code | Meaning | Common cause |
|---|---|---|
| 400 | Malformed body or unsupported parameter | n>1 or response_format on /v1/chat/completions; unsupported audio format; malformed messages |
| 401 | Missing or invalid API key | Wrong header form, revoked key, typo |
| 402 | Insufficient balance | Reservation pre-debit blocked the call before upstream — top up at /dashboard/billing |
| 403 | Email not verified or feature gated | Verify the email link sent at signup |
| 429 | Rate limit / concurrency cap | Max 5 concurrent calls per user; per-route rate limit |
| 500 | Infrastructure error | Always opaque — never leaks upstream error bodies |
| 502 | Upstream returned no usable output | Safety filter blocked image, model 4xx, transient upstream |
| 503 | Transient capacity | Pool exhausted or upstream unavailable — retry with backoff |
| 504 | Long-running operation timed out | Video generation passed the 5-minute ceiling; resubmit |
All error responses follow the SDK shape per route: OpenAI-style { error: { message, type } } on /v1/chat/completions / /v1/images/generations / /v1/audio/speech / /v1/video/generations; Anthropic-style { type: "error", error: { type, message } } on /v1/messages.
Rate limits
- 5 concurrent calls per user across every endpoint. Sixth in-flight request returns 429.
- Pre-call balance reservation: every call pre-debits a worst-case cost before hitting upstream. Insufficient balance returns 402 immediately, no upstream burn.
- Per-route soft rate limits: 60–600 requests/minute depending on route. Backoff in
Retry-Afterwhen you hit one. - We never re-route to a weaker model under load. What you ask for is what runs — or the call returns 429.
Compatibility matrix
What the two surfaces accept side-by-side:
| Feature | /v1/chat/completions | /v1/messages |
|---|---|---|
| Text completion | ✓ | ✓ |
| Streaming | ✓ OpenAI SSE | ✓ Anthropic SSE |
| Multipart content arrays | ✓ full (text + image_url + input_audio + file) | ✓ full |
| Vision (image input) | ✓ (image_url URL or data: URI) | ✓ |
| Audio input | ✓ (input_audio) | ✓ (document/audio block) |
| PDF / document input | ✓ (file part or image_url with .pdf) | ✓ (document block) |
| Tool use / function calling | ✓ (OpenAI tool_calls, every chat model) | ✓ (Anthropic tool_use blocks, every chat model) |
| JSON mode (response_format) | ✗ → 400 | ✗ (planned) |
| n > 1 (multiple completions) | ✗ → 400 | ✗ |
| temperature, top_p, stop | ✓ | ✓ (temperature) |
| Bearer auth | ✓ | ✓ |
| x-api-key auth | ✓ | ✓ |
Support
Open the Discord — a human replies in under an hour, 24/7. Status at lmhaven.xyz/status.