Documentation

LMHaven is a single REST surface for chat, vision, image, voice, and video generation — flat 5× cheaper than first-party list across every model. Two API shapes are supported side-by-side: OpenAI Chat Completions and Anthropic Messages. Drop in by changing one base URL.

Quickstart

  1. Create an API key at dashboard/keys — it starts with sk-lmh-.
  2. Point your favourite SDK at https://lmhaven.xyz/api/v1 as the base URL.
  3. Make a request. Same key works across chat, image, audio, and video.
curl https://lmhaven.xyz/api/v1/chat/completions \
  -H "Authorization: Bearer sk-lmh-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "messages": [{"role":"user","content":"Write a haiku about debugging."}]
  }'

Authentication

Every endpoint accepts the API key on either header — pick whichever your SDK already sends:

  • Authorization: Bearer sk-lmh-… — used by OpenAI SDKs, Cursor, Continue, Cline, Roo Code, and any tool pointed at /api/v1.
  • x-api-key: sk-lmh-… — used by Anthropic SDKs and Claude Code CLI when pointed at /v1/messages.

Keys are workspace-scoped. Revoke from dashboard/keys — subsequent calls fail with 401 immediately, no propagation delay.

Tool integrations

Drop-in configs for the most-asked tools.

Cursor

Settings → Models → OpenAI API Key:

  • API Key: sk-lmh-…
  • Override OpenAI Base URL: https://lmhaven.xyz/api/v1
  • Add custom model: gemini-3-1-pro, glm-5-1, etc.
  • Claude & GPT in Cursor: Cursor special-cases model ids beginning with claude-, gpt-, or o<n> and routes them through its own provider logic (ignoring this base URL), so they fail. Use the lmh--prefixed aliases instead, which Cursor treats as plain custom models: lmh-claude-opus-4-7, lmh-claude-sonnet-4-6, lmh-gpt-5-5, lmh-gpt-5-2, lmh-o4-mini. Same models, same pricing.

Continue.dev

Continue uses YAML (~/.continue/config.yaml) now — the older config.json format is deprecated. Use provider: openai for every model (chat models route through the OpenAI-compatible endpoint, including Claude).

~/.continue/config.yaml
name: LMHaven
version: 0.0.1
schema: v1
models:
  - name: Claude Opus 4.7
    provider: openai
    model: claude-opus-4-7
    apiBase: https://lmhaven.xyz/api/v1
    apiKey: sk-lmh-…
    roles:
      - chat
      - edit
      - apply
  - name: Kimi K2.6
    provider: openai
    model: kimi-k2-6
    apiBase: https://lmhaven.xyz/api/v1
    apiKey: sk-lmh-…
    roles:
      - chat
      - edit

Cline

Cline → Settings → API Provider: OpenAI Compatible

  • Base URL: https://lmhaven.xyz/api/v1
  • API Key: sk-lmh-…
  • Model: any from the catalog (claude-sonnet-4-6 recommended).

Roo Code

Same as Cline. Provider type: OpenAI Compatible.

Aider

Aider supports any OpenAI-compatible endpoint. Two env vars + the model flag:

bash
export OPENAI_API_BASE="https://lmhaven.xyz/api/v1"
export OPENAI_API_KEY="sk-lmh-…"

aider --model openai/claude-sonnet-4-6   # or any chat model id from /v1/models

Droid (Factory.ai)

In Droid's BYOK config, add each model with:

  • Provider type: generic-chat-completion-api (or openai)
  • Base URL: https://lmhaven.xyz/api
  • API key: your sk-lmh-…
  • Model name: any chat model id from /v1/models

Tool use works on every chat model — Droid will see tool_calls the same way it does on OpenAI directly. No need to set provider: "anthropic" for Claude / Gemini models; one config covers every chat model.

Claude Code CLI

Claude Code speaks the Anthropic Messages API. Claude Code appends /v1/messages to the base URL, so point the base at /api (it resolves to /api/v1/messages). Two env vars:

bash
export ANTHROPIC_BASE_URL="https://lmhaven.xyz/api"
export ANTHROPIC_AUTH_TOKEN="sk-lmh-…"
# Some Claude Code builds also read ANTHROPIC_API_KEY:
export ANTHROPIC_API_KEY="sk-lmh-…"

claude  # any claude-* model id resolves automatically (opus / sonnet / haiku tier)

API reference

GET /v1/models

OpenAI-shaped catalog discovery. Public — no auth needed.

bash
curl https://lmhaven.xyz/api/v1/models
json
{
  "object": "list",
  "data": [
    { "id": "claude-opus-4-7", "object": "model", "owned_by": "lmhaven", … },
    { "id": "gemini-3-1-pro", … }
  ]
}

POST /v1/chat/completions

OpenAI-compatible. Same shape as platform.openai.com/docs/api-reference/chat/create.

Supported fields: model, messages, max_tokens, system, stream, temperature, top_p, stop (string or array, capped at 5), tools, tool_choice, reasoning_effort, thinking_budget.

Reasoning control for thinking-capable models (Gemini 2.5+/3.x, GLM, etc.). Set reasoning_effort to one of "none", "low", "medium", "high" for preset budgets, or pass an explicit integer via thinking_budget when you need fine-grained control (Antigravity-style). For Gemini the preset map is none=0, low=128, medium=2048, high=16384; Gemini Pro clamps to its 128 minimum when thinking is enabled. OpenAI's "minimal" alias is accepted and treated as "none".

Multimodal content arrays ([{type:"text",…}, {type:"image_url",…}, {type:"input_audio",…}]) are fully supported. Pass image_url.url as either an https://… URL (we fetch and inline server-side, SSRF-hardened) or a data:image/…;base64,… URI (parsed locally). PDFs and video go through the same field — infer-by-extension picks the right size cap and MIME allowlist.

Tool use: tools, tool_choice, role:"tool" messages, and assistant.tool_calls all work across every chat model — Claude, GPT, Gemini, GLM, Grok, DeepSeek, Qwen, Kimi, Llama, Mistral. Streaming emits OpenAI-shape tool_calls deltas. System prompts work either via top-level system or as a role:"system" message — both paths are normalized.

Rejected with 400 (loud, not silent): response_format (json_mode / structured outputs), n > 1.

curl https://lmhaven.xyz/api/v1/chat/completions \
  -H "Authorization: Bearer sk-lmh-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "messages": [
      {"role": "system", "content": "You are a precise SQL tutor."},
      {"role": "user", "content": "Explain window functions in two sentences."}
    ],
    "stream": true,
    "temperature": 0.4
  }'

POST /v1/messages

Anthropic-compatible. Same shape as docs.anthropic.com/en/api/messages. Use this for vision, tool use, and any Claude SDK consumer (including Claude Code CLI).

Full feature parity: messages (string or content blocks), system (string or text-block array), max_tokens, temperature, stream, tools, tool_choice, thinking. Image content blocks ({type:"image",…}) and tool_use / tool_result blocks all flow through.

Extended thinking in Anthropic shape — pass "thinking": { "type": "enabled", "budget_tokens": 4096 } to enable, or "thinking": { "type": "disabled" } / null to turn it off. When the request routes to Gemini Pro this is mapped to thinkingConfig.thinkingBudget with a floor of 128. For Claude SKUs the field is forwarded verbatim to the upstream Anthropic-shape API.

Model aliases: versioned Anthropic ids are accepted by tier — claude-sonnet-4-5, claude-3-5-sonnet-20241022, claude-haiku-4-5-latest all resolve to the matching tier in our catalog (claude-opus-4-7, claude-sonnet-4-6, claude-haiku-4-5), so SDKs that pin to a specific Anthropic version work out of the box.

curl https://lmhaven.xyz/api/v1/messages \
  -H "x-api-key: sk-lmh-…" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "max_tokens": 1024,
    "messages": [{"role":"user","content":"hi"}]
  }'

POST /v1/images/generations

OpenAI-shaped. Returns 24-hour presigned URLs.

Supported: model, prompt, size, aspect_ratio, images (1–10 reference images for editing / composition), n (1–4 outputs per call; each billed at the full per-image rate).
Rejected with 400: quality, style, response_format != "url".

bash
curl https://lmhaven.xyz/api/v1/images/generations \
  -H "Authorization: Bearer sk-lmh-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nano-banana-2",
    "prompt": "A watercolor city skyline at dusk, cinematic lighting",
    "aspect_ratio": "16:9",
    "n": 4
  }'

With reference images. Pass images as an array of data URLs (or raw base64) — they become inline-image parts on the generation request for editing / composition. Up to 10 reference images per call.

bash
curl https://lmhaven.xyz/api/v1/images/generations \
  -H "Authorization: Bearer sk-lmh-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nano-banana-2",
    "prompt": "Place this character on a snowy mountain at golden hour.",
    "aspect_ratio": "16:9",
    "images": [
      "data:image/png;base64,iVBORw0KGgoAAAANSUhEUg…",
      "data:image/jpeg;base64,/9j/4AAQSkZJRgABAQEAYABg…"
    ],
    "n": 2
  }'

POST /v1/audio/speech

OpenAI-shaped TTS. Returns audio bytes inline.

Voice aliases (OpenAI names): alloy · echo · fable · onyx · nova · shimmer · coral · ash · ballad · sage · verse — each maps to the closest stock ElevenLabs voice we host.

ElevenLabs voice names: rachel · adam · bella · antoni · domi · elli · josh · arnold · sam. Or pass any raw ElevenLabs voice id directly.

Formats: mp3 (default), pcm, wav. opus/aac/flac rejected with 400.

bash
curl https://lmhaven.xyz/api/v1/audio/speech \
  -H "Authorization: Bearer sk-lmh-…" \
  -H "Content-Type: application/json" \
  -o output.mp3 \
  -d '{
    "model": "elevenlabs-v3-multilingual",
    "voice": "alloy",
    "input": "Hello from lmhaven."
  }'

POST /v1/video/generations

Submit prompt → block until ready → 24-hour presigned URLs. Veo 3 / Veo 3 Fast emit native audio in the clip. Generation takes 30–90s typical, 5-minute hard ceiling.

Supported: model, prompt, ratio, duration, resolution, image (first-frame reference, Veo only), end_image (last-frame reference for frames-to-video transition), n (1–4 videos per call; each billed at the full per-video rate).

bash
curl https://lmhaven.xyz/api/v1/video/generations \
  -H "Authorization: Bearer sk-lmh-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "veo-3.1",
    "prompt": "A calm ocean wave at sunset, slow camera pan",
    "ratio": "16:9",
    "duration": 8,
    "resolution": "720p",
    "n": 1
  }'

Image-to-video. Pass image as a data URL — Veo animates from that still as the first frame. Add end_image to lock the last frame too (Veo generates a transition between the two stills).

bash
curl https://lmhaven.xyz/api/v1/video/generations \
  -H "Authorization: Bearer sk-lmh-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "veo-3.1",
    "prompt": "Subject walks toward the camera, slow zoom in",
    "ratio": "16:9",
    "duration": 8,
    "image": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUg…",
    "end_image": "data:image/png;base64,iVBORw0KGgoAAAANSUhEUg…"
  }'
json
{
  "data": [
    { "url": "https://cdn.lmhaven.xyz/videos/…/abc.mp4?…", "duration_seconds": 8 },
    { "url": "https://cdn.lmhaven.xyz/videos/…/def.mp4?…", "duration_seconds": 8 }
  ]
}

POST /v1/files

Upload media once, reference it from many chat calls by file_id. The OpenAI and Anthropic SDKs both call this path automatically when you pass a File handle to a chat method — drop-in for either ecosystem. Multipart file + purpose body.

Limits: 200 files per account, 2 GB total, 30-day TTL (lazy-expired on read past TTL). Per-file caps mirror inline upload — 20 MB image / video / audio, 32 MB PDF. Same MIME allowlist applies.

curl https://lmhaven.xyz/api/v1/files \
  -H "Authorization: Bearer sk-lmh-…" \
  -F purpose=vision \
  -F file=@/path/to/photo.png

Models & pricing

Flat 5× cheaper than first-party across every model — no tiered SKUs, no cache discount math by default. Rates verified weekly against each provider’s published list price (sources cited inline in lib/models.ts).

Text · vision (per 1M tokens, input / output)

ModelLMHavenFirst-party
claude-fable-5$2.00/1M in · $10.00/1M outAnthropic $10.00/1M / $50.00/1M
claude-opus-4-7$1.00/1M in · $5.00/1M outAnthropic $5.00/1M / $25.00/1M
claude-opus-4-8$1.00/1M in · $5.00/1M outAnthropic $5.00/1M / $25.00/1M
claude-opus-5$1.00/1M in · $5.00/1M outAnthropic $5.00/1M / $25.00/1M
claude-sonnet-5$0.60/1M in · $3.00/1M outAnthropic $3.00/1M / $15.00/1M
claude-sonnet-4-6$0.60/1M in · $3.00/1M outAnthropic $3.00/1M / $15.00/1M
claude-haiku-4-5$0.20/1M in · $1.00/1M outAnthropic $1.00/1M / $5.00/1M
gpt-5-2$1.00/1M in · $8.00/1M outOpenAI $5.00/1M / $40.00/1M
gpt-5-4$1.00/1M in · $8.00/1M outOpenAI $5.00/1M / $40.00/1M
gpt-5-5$1.25/1M in · $10.00/1M outOpenAI $6.25/1M / $50.00/1M
gpt-5-6-sol$1.00/1M in · $6.00/1M outOpenAI $5.00/1M / $30.00/1M
gpt-5-6-terra$0.50/1M in · $3.00/1M outOpenAI $2.50/1M / $15.00/1M
gpt-5-6-luna$0.20/1M in · $1.20/1M outOpenAI $1.00/1M / $6.00/1M
gpt-5-mini$0.05/1M in · $0.40/1M outOpenAI $0.25/1M / $2.00/1M
o4-mini$0.22/1M in · $0.88/1M outOpenAI $1.10/1M / $4.40/1M
gemini-3-1-pro$0.50/1M in · $4.00/1M outGoogle $2.50/1M / $20.00/1M
gemini-3-5-flash$0.06/1M in · $0.50/1M outGoogle $0.30/1M / $2.50/1M
gemini-3-1-flash-lite$0.02/1M in · $0.08/1M outGoogle $0.10/1M / $0.40/1M
llama-4-maverick$0.07/1M in · $0.28/1M outMeta $0.35/1M / $1.40/1M
llama-4-scout$0.04/1M in · $0.16/1M outMeta $0.20/1M / $0.80/1M
deepseek-r1$0.11/1M in · $0.44/1M outDeepSeek $0.55/1M / $2.19/1M
deepseek-v3-2$0.05/1M in · $0.22/1M outDeepSeek $0.27/1M / $1.10/1M
qwen-3-coder$0.20/1M in · $0.90/1M outAlibaba $1.00/1M / $4.50/1M
qwen-3-next$0.03/1M in · $0.12/1M outAlibaba $0.15/1M / $0.60/1M
kimi-k2-thinking$0.12/1M in · $0.50/1M outMoonshot $0.60/1M / $2.50/1M
kimi-k2-6$0.15/1M in · $0.60/1M outMoonshot $0.75/1M / $3.00/1M
kimi-k2-7-code$0.19/1M in · $0.80/1M outMoonshot $0.95/1M / $4.00/1M
deepseek-v4-pro$0.10/1M in · $0.50/1M outDeepSeek $0.50/1M / $2.50/1M
deepseek-v4-flash$0.04/1M in · $0.16/1M outDeepSeek $0.20/1M / $0.80/1M
grok-4-3$0.80/1M in · $4.00/1M outxAI $4.00/1M / $20.00/1M
gpt-oss-120b$0.05/1M in · $0.20/1M outOpenAI $0.25/1M / $1.00/1M
grok-4-1-fast$0.20/1M in · $1.00/1M outxAI $1.00/1M / $5.00/1M
grok-4-2$0.60/1M in · $3.00/1M outxAI $3.00/1M / $15.00/1M
codestral-2$0.06/1M in · $0.18/1M outMistral $0.30/1M / $0.90/1M
mistral-medium-3$0.08/1M in · $0.40/1M outMistral $0.40/1M / $2.00/1M

Image (per call)

ModelLMHavenFirst-party
nano-banana-pro$0.008Google $0.039
nano-banana-2$0.027Google $0.134

Voice (per 1K characters)

ModelLMHavenFirst-party

Video (per 8s clip)

ModelLMHavenFirst-party
veo-3.1$0.640Google $3.200
veo-3.1-fast$0.160Google $0.800
veo-3$0.640Google $3.200
veo-3-fast$0.160Google $0.800
veo-2$0.800Google $4.000

Reseller / enterprise rates: per-account discount factors are available (8×, 10×, 20× cheaper than list). Contact us via Discord for contracted tiers.

Multimodal input

Both endpoints accept image, video, audio, and PDF input on every vision-capable model (claude-*, gemini-3-*). Two ref forms, both fully supported:

  • Inline base64 — Anthropic {type:"image", source:{type:"base64",media_type,data}} or OpenAI data:image/png;base64,… URIs. Parsed locally.
  • URL (https only) — Anthropic {type:"image", source:{type:"url",url}} or OpenAI image_url.url: "https://…". We fetch the URL server-side with an SSRF guard (RFC1918 / link-local / cloud-metadata addresses are pre-rejected by DNS resolution; redirects re-validated per hop, capped at 3) and inline as base64 before forwarding.

Limits: 20 MB image / video / audio, 32 MB PDF, 10 s fetch timeout. Allowed MIMEs: PNG / JPEG / GIF / WEBP / HEIC / HEIF for image; MP4 / WebM / MPEG / MOV / AVI / 3GPP for video; MP3 / WAV / AAC / M4A / OGG / FLAC for audio; PDF for documents. Anthropic document blocks (PDF + plain-text + structured-content forms) are accepted on /v1/messages.

Files API references are also live — upload once via POST /v1/files, then pass the resulting file_id as {type:"file", file:{file_id}} on /v1/chat/completions or as {type:"image", source:{type:"file", file_id}} on /v1/messages. Per-account isolation is enforced server-side — a file_id from another account returns 404.

Anthropic shape — image (URL)

curl https://lmhaven.xyz/api/v1/messages \
  -H "x-api-key: sk-lmh-…" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "max_tokens": 512,
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "What is in this image?"},
        {
          "type": "image",
          "source": {
            "type": "url",
            "url": "https://example.com/photo.png"
          }
        }
      ]
    }]
  }'

Anthropic shape — image (base64)

# base64-encode locally first
B64=$(base64 -i photo.png)
curl https://lmhaven.xyz/api/v1/messages \
  -H "x-api-key: sk-lmh-…" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d "{
    \"model\": \"claude-sonnet-4-6\",
    \"max_tokens\": 512,
    \"messages\": [{
      \"role\": \"user\",
      \"content\": [
        {\"type\": \"text\", \"text\": \"What is in this image?\"},
        {
          \"type\": \"image\",
          \"source\": {
            \"type\": \"base64\",
            \"media_type\": \"image/png\",
            \"data\": \"$B64\"
          }
        }
      ]
    }]
  }"

Anthropic shape — PDF document

curl https://lmhaven.xyz/api/v1/messages \
  -H "x-api-key: sk-lmh-…" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-opus-4-7",
    "max_tokens": 1024,
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Summarize this contract in three bullets."},
        {
          "type": "document",
          "source": {"type": "url", "url": "https://example.com/contract.pdf"},
          "title": "Master Services Agreement",
          "context": "Reviewing for renewal terms"
        }
      ]
    }]
  }'

OpenAI shape — image_url + input_audio

curl https://lmhaven.xyz/api/v1/chat/completions \
  -H "Authorization: Bearer sk-lmh-…" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-4-6",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "Describe this image."},
        {"type": "image_url", "image_url": {"url": "https://example.com/photo.png"}}
      ]
    }]
  }'

Files API — upload once, reuse by file_id

# Returns {"id": "file_…", "object": "file", "bytes": …, "filename": …, "mime_type": …}
curl https://lmhaven.xyz/api/v1/files \
  -H "Authorization: Bearer sk-lmh-…" \
  -F purpose=vision \
  -F file=@photo.png

Tool use (function calling)

Works on both endpoints, across every chat model in the catalog — Claude, GPT, Gemini, GLM, Grok, DeepSeek, Qwen, Kimi, Llama. No per-model configuration needed. Whichever shape your client speaks, point it at LMHaven and tools just work.

  • OpenAI shape on /v1/chat/completions — tools, tool_choice, role:"tool", assistant.tool_calls. Streaming emits tool_calls deltas. Used by Cursor, Continue, Cline, Roo Code, Aider, Droid, OpenAI SDKs.
  • Anthropic shape on /v1/messages — tools array, tool_use response blocks, tool_result in user follow-ups, input_json_delta streaming. Used by Claude Code CLI, Anthropic SDKs.

Translation between OpenAI and Anthropic tool shapes happens server-side and is transparent — your client sees the shape it asked for, no matter which model serves the request.

Streaming

OpenAI shape (/v1/chat/completions)

SSE with data: {chat.completion.chunk} frames, terminated by data: [DONE]. Each chunk carries choices[0].delta.content for text, or choices[0].delta.tool_calls for function-call deltas. Final chunk has choices[0].finish_reason set to "stop" for plain text or "tool_calls" when the model is asking to call a function, plus a usage object.

Anthropic shape (/v1/messages)

Full Anthropic event sequence: message_start → content_block_start → content_block_delta×N → content_block_stop → message_delta → message_stop. Tool-use blocks emit input_json_delta with the full args.

Prompt caching

Pass cache_control headers in your request — upstream honours them for latency wins regardless of any LMHaven setting.

Cache discounts on the bill are an opt-in toggle on /dashboard/keys. With it enabled:

  • Anthropic cache reads: 10% of input rate (~50× off list net).
  • Anthropic cache writes: 125% of input rate (one-time premium per cached prefix; break-even after ~2 reads).
  • Gemini cache reads: 25% of input rate (~20× off list net).
  • Gemini cache writes: 100% of input rate (no premium).
  • Output tokens: unaffected, always at your standard rate.

Quality is identical with the toggle on or off. Caching only changes billing — the model still runs full inference on every call. Caching reuse is at the prefix-processing layer, not response replay.

Errors

CodeMeaningCommon cause
400Malformed body or unsupported parametern>1 or response_format on /v1/chat/completions; unsupported audio format; malformed messages
401Missing or invalid API keyWrong header form, revoked key, typo
402Insufficient balanceReservation pre-debit blocked the call before upstream — top up at /dashboard/billing
403Email not verified or feature gatedVerify the email link sent at signup
429Rate limit / concurrency capMax 5 concurrent calls per user; per-route rate limit
500Infrastructure errorAlways opaque — never leaks upstream error bodies
502Upstream returned no usable outputSafety filter blocked image, model 4xx, transient upstream
503Transient capacityPool exhausted or upstream unavailable — retry with backoff
504Long-running operation timed outVideo generation passed the 5-minute ceiling; resubmit

All error responses follow the SDK shape per route: OpenAI-style { error: { message, type } } on /v1/chat/completions / /v1/images/generations / /v1/audio/speech / /v1/video/generations; Anthropic-style { type: "error", error: { type, message } } on /v1/messages.

Rate limits

  • 5 concurrent calls per user across every endpoint. Sixth in-flight request returns 429.
  • Pre-call balance reservation: every call pre-debits a worst-case cost before hitting upstream. Insufficient balance returns 402 immediately, no upstream burn.
  • Per-route soft rate limits: 60–600 requests/minute depending on route. Backoff in Retry-After when you hit one.
  • We never re-route to a weaker model under load. What you ask for is what runs — or the call returns 429.

Compatibility matrix

What the two surfaces accept side-by-side:

Feature/v1/chat/completions/v1/messages
Text completion✓✓
Streaming✓ OpenAI SSE✓ Anthropic SSE
Multipart content arrays✓ full (text + image_url + input_audio + file)✓ full
Vision (image input)✓ (image_url URL or data: URI)✓
Audio input✓ (input_audio)✓ (document/audio block)
PDF / document input✓ (file part or image_url with .pdf)✓ (document block)
Tool use / function calling✓ (OpenAI tool_calls, every chat model)✓ (Anthropic tool_use blocks, every chat model)
JSON mode (response_format)✗ → 400✗ (planned)
n > 1 (multiple completions)✗ → 400✗
temperature, top_p, stop✓✓ (temperature)
Bearer auth✓✓
x-api-key auth✓✓

Support

Open the Discord — a human replies in under an hour, 24/7. Status at lmhaven.xyz/status.

Docs · LMHaven