API reference
Integrate once at https://api.zenllm.org. Use the same API key for the catalog, OpenAI-compatible requests, and Anthropic Messages.
Introduction
ZenLLM is one API in front of every frontier model, with automatic provider failover and prepaid per-token billing. The base URL for everything is https://api.zenllm.org.
Chat routes are wire-compatible with the OpenAI and Anthropic SDKs: point your existing client at the base URL and keep your code. Jev uses the separate TypeSafe System One endpoint. Streaming uses standard SSE on chat models, tool calling passes through natively, and responses always echo the model id you asked for, whichever upstream served it.
curl https://api.zenllm.org/v1/chat/completions \
-H "Authorization: Bearer $ZENLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-fable-5.1",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'Authentication
Create a key in the dashboard. Keys are shown once and stored hashed; each key can carry its own model allowlist, rate limit, IP allowlist, total spend cap and expiry.
Send it on any endpoint using whichever header your SDK already uses.
# Any one of these works on every endpoint
Authorization: Bearer sk-zenllm-...
x-api-key: sk-zenllm-...
anthropic-api-key: sk-zenllm-... # for Anthropic SDKsErrors and limits
Errors are JSON with a stable shape and never leak upstream provider details. Account rate limits scale with your balance plus lifetime spend (spending credits never lowers your tier): $5+ gives 15 rpm, $10+ 30 rpm, $50+ 70 rpm, $100+ 135 rpm and $200+ 250 rpm; contact support for more.
Failed requests are refunded automatically: billing settles only on completed responses.
- 401auth_error
- Missing or invalid API key, or an IP outside the key's allowlist.
- 402billing_error
- Balance too low, or a key/account spend limit was reached.
- 404not_found
- Unknown model id or route.
- 429rate_limit_error
- Key RPM or account tier exceeded. Honor Retry-After with jitter.
- 502upstream_error
- Every provider in the model's chain failed. Safe to retry.
{
"error": {
"message": "Insufficient credits",
"type": "billing_error",
"code": 402
}
}Overview
Anything that speaks the OpenAI or Anthropic APIs works unchanged, so every tool below connects with just a base URL and a key. Any model id from /v1/models works in any tool: run Claude in an OpenAI client or GPT in an Anthropic one, the router translates.
CORS is open on the /v1 surface, so browser-based tools can call the API directly with your key; no proxy needed. Create your key in the dashboard first, it is shown once.
SillyTavern
The most common RP frontend. Connects through its Chat Completion API.
- 1Open the plug icon (API Connections).
- 2Set API to
Chat Completionand the source toCustom (OpenAI-compatible). - 3Paste the endpoint
https://api.zenllm.org/v1and your key, then press Connect. - 4Pick any model from the list; it loads straight from the live catalog.
Presets work as-is, including prompt post-processing, system prompts and assistant prefill.
API: Chat Completion
Chat Completion Source: Custom (OpenAI-compatible)
Custom Endpoint (Base URL): https://api.zenllm.org/v1
API Key: sk-zenllm-...
Model: claude-fable-5.1 # list fills from /v1/modelsJanitorAI
JanitorAI calls the API from your browser, which works because CORS is open on /v1. Your key stays in your browser's local storage.
- 1Open a chat and go to API Settings.
- 2Choose
Proxy, then configurationCustom. - 3Set the proxy URL to
https://api.zenllm.org/v1/chat/completions(the full path, not just the base). - 4Enter your key and a model name, save, then use "Check proxy" to verify.
API: Proxy
Configuration: Custom
Model name: claude-fable-5.1
Proxy URL: https://api.zenllm.org/v1/chat/completions
API Key: sk-zenllm-...Marinara Engine
The local chat, roleplay and game engine. Its custom provider speaks the OpenAI format, and connections can be overridden per chat, so different characters can run on different models.
- 1Open Connections and add a new connection.
- 2Choose
Custom (OpenAI-compatible)as the provider. - 3Set the base URL to
https://api.zenllm.org/v1and paste your key. - 4The model picker fills itself from the catalog; pick one and save.
Connections -> Add Connection
Provider: Custom (OpenAI-compatible)
Base URL: https://api.zenllm.org/v1
API Key: sk-zenllm-...
# the model picker fills itself from /v1/modelsAgnai and Risu
Both connect through their OpenAI-compatible option. Agnai takes the base URL; Risu wants the full chat completions path.
- 1Agnai: in your preset, pick the OpenAI-compatible (third party) service and set the API URL to
https://api.zenllm.org/v1. - 2Risu: under Settings, choose the custom OpenAI-compatible model and set the request URL to
https://api.zenllm.org/v1/chat/completions. - 3Paste your key and type any model id from the catalog.
Preset -> AI Service: OpenAI-compatible / third party
API URL: https://api.zenllm.org/v1
API Key: sk-zenllm-...
Model: claude-fable-5.1Claude Code
Claude Code speaks the Anthropic Messages API, which ZenLLM serves natively at /v1/messages, streaming, tool use and token counting included.
- 1Export
ANTHROPIC_BASE_URLandANTHROPIC_AUTH_TOKENin your shell, or put them in~/.claude/settings.jsonto make it permanent. - 2Optionally pin
ANTHROPIC_MODELto any catalog id. - 3Run
claude. No login needed; the key is the auth.
export ANTHROPIC_BASE_URL="https://api.zenllm.org"
export ANTHROPIC_AUTH_TOKEN="sk-zenllm-..."
export ANTHROPIC_MODEL="claude-fable-5.1"
claudeCodex CLI
Codex talks the Responses API, served at /v1/responses. Point it at ZenLLM with a custom model provider.
- 1Add the provider block to
~/.codex/config.toml. - 2Export
ZENLLM_API_KEYin your shell. - 3Run
codex. Switch models any time with/modelor themodelline in the config.
# ~/.codex/config.toml
model = "gpt-6-astra"
model_provider = "zenllm"
[model_providers.zenllm]
name = "ZenLLM"
base_url = "https://api.zenllm.org/v1"
env_key = "ZENLLM_API_KEY"
wire_api = "responses"opencode
Register ZenLLM as a provider through the OpenAI-compatible SDK adapter.
- 1Add the provider block to
~/.config/opencode/opencode.json(or a project-localopencode.json). - 2Export
ZENLLM_API_KEY, or paste the key inline if you prefer. - 3Run
opencodeand pick a model with/models.
// ~/.config/opencode/opencode.json
{
"provider": {
"zenllm": {
"npm": "@ai-sdk/openai-compatible",
"name": "ZenLLM",
"options": {
"baseURL": "https://api.zenllm.org/v1",
"apiKey": "{env:ZENLLM_API_KEY}"
},
"models": {
"claude-fable-5.1": { "name": "Claude Fable 5.1" },
"gpt-6-astra": { "name": "GPT-6 Astra" }
}
}
}
}Zed
Zed's agent panel supports OpenAI-compatible providers natively.
- 1Add the provider under
language_models.openai_compatiblein yoursettings.json, listing the models you want. - 2Open the agent panel's settings, find ZenLLM and paste your key when prompted.
- 3Select the model from the panel's model picker.
// Zed settings.json
{
"language_models": {
"openai_compatible": {
"ZenLLM": {
"api_url": "https://api.zenllm.org/v1",
"available_models": [{
"name": "claude-fable-5.1",
"display_name": "Claude Fable 5.1",
"max_tokens": 1000000
}]
}
}
}
}Cursor
Cursor can route its chat models through any OpenAI-compatible endpoint via the base URL override.
- 1Open Cursor Settings, then Models.
- 2Paste your ZenLLM key into the OpenAI API key field and enable the base URL override with
https://api.zenllm.org/v1. - 3Add the catalog ids you want as custom model names, then hit Verify.
Note: the override applies to chat. Cursor-native features like Tab completions still run on Cursor's own service.
Cursor Settings -> Models
OpenAI API Key: sk-zenllm-...
Override OpenAI Base URL: https://api.zenllm.org/v1
Model names: claude-fable-5.1, gpt-6-astra, ...Cline and Continue
The two big VS Code agents, both with first-class OpenAI-compatible support.
- 1Cline: in provider settings choose
OpenAI Compatible, set the base URL, key and a model id. - 2Continue: add a model entry with
provider: openaiandapiBaseto~/.continue/config.yaml.
API Provider: OpenAI Compatible
Base URL: https://api.zenllm.org/v1
API Key: sk-zenllm-...
Model ID: claude-fable-5.1LibreChat and Open WebUI
Self-hosted chat UIs. Both can fetch the model list from the catalog, so new models appear without config changes.
- 1LibreChat: add a custom endpoint to
librechat.yamlwithfetch: trueand restart. - 2Open WebUI: in Admin Panel connections, enable the OpenAI API and set the base URL and key.
# librechat.yaml
version: 1.3.13
cache: true
endpoints:
custom:
- name: 'ZenLLM'
apiKey: '${ZENLLM_API_KEY}'
baseURL: 'https://api.zenllm.org/v1'
models:
default:
- 'gpt-6-sol'
- 'gpt-6-luna'
- 'kimi-k3'
- 'claude-fable-5.1'
fetch: true
titleConvo: true
titleModel: 'gpt-6-luna'
modelDisplayLabel: 'ZenLLM'
customParams:
defaultParamsEndpoint: 'openAI'
reasoningFormat: 'reasoning_effort'
reasoningKey: 'reasoning_content'
includeReasoningContent: trueMissing or invalid API key
The request had no key, a malformed key, or a key that does not exist. Keys are sent per request; there are no sessions on the API surface.
Common causes, in the order we see them:
- 1The header is missing or misspelled. Any of
Authorization: Bearer,x-api-keyoranthropic-api-keyworks. - 2The key was deleted or has expired (check its expiry in the dashboard).
- 3The key is frozen, or the request came from an IP outside the key's allowlist.
- 4Whitespace or quotes were pasted along with the key.
Verify a key without spending credits: call GET /v1/keyinfo with the key. Repeated failures from one IP are throttled, so scripts with a bad key will start seeing errors quickly.
HTTP/2 401
{
"error": {
"message": "Invalid API key. Pass it as 'Authorization: Bearer <key>'.",
"type": "auth_error",
"code": 401
}
}Balance too low
Every request reserves an estimated cost before running and settles to the exact cost after. A 402 means the reservation could not be covered: your prepaid balance is lower than the estimated cost of this request.
Large max_tokens values raise the estimate, so a request can 402 even when your balance covers the typical response. Lower max_tokens or top up in the dashboard. Failed requests are never charged, and mid-stream cancels only pay for tokens already generated.
HTTP/2 402
{
"error": {
"message": "Insufficient credits",
"type": "billing_error",
"code": 402
}
}Spending limit exhausted
Keys and accounts can carry a total spend cap. Once the lifetime spend of a key (or the whole account) reaches its cap, requests return 402 even when the balance is positive. This is the guard rail working, not a billing failure.
- 1Check
GET /v1/keyinfo: it returns the key'sspendLimitUsdnext to itsspentUsd. - 2Raise or clear the cap in the dashboard (API keys for per-key caps, Settings for the account cap), or rotate to a fresh key.
Unknown model or route
The model id is not in the catalog, or the URL path does not exist. Model ids are exact strings from GET /v1/models; there are no version wildcards and no legacy aliases. If your client is configured with a name like gpt-4o, switch it to a real catalog id.
For routes: the base URL is https://api.zenllm.org and endpoints live under /v1/.... If your client appends /v1 itself, do not include it in the base URL twice.
HTTP/2 404
{
"error": {
"message": "model not found: gpt-9000",
"type": "not_found",
"code": 404
}
}Model refusal
The model's own safety layer blocked the completion before any output was generated. This is the model provider's safeguard, not a ZenLLM policy: the same prompt sent to the lab directly would be blocked the same way.
Retrying the identical request will not help; the decision is deterministic for a given prompt. Rephrase the request or switch models. When a model declines in plain text instead ("I can't help with that"), that is a normal 200 completion, not this error.
HTTP/2 403
{
"error": {
"message": "The model is refusing to complete the request (Safeguards).",
"type": "model_refusal",
"code": "model_refusal"
}
}Rate limit exceeded
Two independent limits produce a 429, and either can fire:
- 1Key RPM: each key carries its own requests-per-minute setting (dashboard, per key). Good for capping one app.
- 2Account tier: your account-wide RPM scales with balance plus lifetime spend: $5+ gives 15 rpm, $10+ 30, $50+ 70, $100+ 135, $200+ 250. Spending never lowers your tier.
Honor the retry-after header with jitter instead of hammering: retries that ignore it count against the same window. Your current limits are visible on GET /v1/keyinfo (rpmLimit for the key, tierRpm for the account).
HTTP/2 429
retry-after: 21
{
"error": {
"message": "Rate limit exceeded",
"type": "rate_limit_error",
"code": 429
}
}Invalid request
The request itself is malformed: an unknown role, a prompt over the model's context window, an unsupported parameter value, or an image sent to a text-only model. The message carries the specific reason, so surface it to your users; a retry with the same body will fail the same way.
Nothing is charged for a 400. Context-window overflows are best fixed by trimming history or switching to a 1M-context model (most of the catalog).
HTTP/2 400
{
"error": {
"message": "reasoning effort must be one of low, medium, high, xhigh, max (got 'turbo')",
"type": "invalid_request_error",
"code": 400
}
}Every provider failed
Each model runs on an ordered chain of upstream providers, and a request only fails after the whole chain has been tried. A 502 therefore means every provider for that model errored: an upstream outage, not something wrong with your request.
- 1Retry with backoff; failures are usually seconds-long blips and requests that fail cost nothing.
- 2If it persists, switch models: outages rarely span families.
- 3Check
/healthand the Discord for incident notes if a whole family stays down.
HTTP/2 502
{
"error": {
"message": "The upstream model provider failed. Try again or use another model.",
"type": "server_error",
"code": 502
}
}Prompt caching
Repeated prompt prefixes are cached upstream and re-read at a fraction of the input price (up to 97% off on some models; see each model's cached_input rate in /v1/models). For Claude and Qwen models the router manages this automatically: on every multi-turn conversation it pins a cache breakpoint on your latest message, so each turn re-reads the whole prior conversation at the cached rate and only the new tokens bill at full price.
Want manual control? Place OpenRouter-style cache_control markers on any text content part, like {"type": "text", "text": "...", "cache_control": {"type": "ephemeral"}} on /v1/chat/completions, /v1/messages or /v1/responses. When you place your own markers the router never adds or moves them.
GPT models cache automatically with no markers; the router adds a stable per-account routing key and 24-hour cache retention for you, and you can pass your own prompt_cache_key to override.
Cache writes are billed where the upstream bills them (for example 1.25x the input rate on Claude models) and itemized on your ledger receipts. Each model's cache_write rate is in /v1/models; models with no write premium bill written tokens at the plain input rate.
Web search
Models with web_search in their supported_features (currently the Claude Fable family) can search the web server-side. Enable it by sending "web_search_options": {} or a {"type": "web_search"} tool on chat completions, or the native Anthropic web_search_20260209 server tool on /v1/messages.
Searches cost a flat $0.01 each (the official rate, shown as web_search_per_call in /v1/models), billed on top of tokens and itemized in your ledger. Results come back with URL citations in the response.
Personal pricing
Accounts can carry custom per-model discounts. If yours does, /v1/models shows your rates whenever you call it with your API key (anonymous calls show list prices), your billing page lists the active deals, and every ledger receipt reproduces the exact rates you were charged.
Cost in every response
Every response tells you what it cost. The usage object carries cost (a number, in dollars) and cost_display (the same amount as readable text), on every endpoint and in both streaming and non-streaming responses.
Non-streaming responses report the exact amount your balance was charged. Streams stamp the running cost into each usage frame, so the final one carries the total and it matches your ledger receipt for the request.
"usage": {
"prompt_tokens": 1204,
"completion_tokens": 388,
"total_tokens": 1592,
"cost": 0.000354,
"cost_display": "$0.000354"
}Chat completions
The OpenAI-compatible endpoint and the one most SDKs use. On failure the router retries the model's next provider inside the same request, so a single call survives a provider outage.
Billing settles on completion: with stream: true the final SSE chunk before [DONE] carries the usage object, including cost and cost_display, and that is what your balance is charged. If you cancel mid-stream you pay only for the tokens generated up to the cancel; requests that fail outright cost nothing. Every response carries an x-zenllm-request-id header you can quote to support and match against your ledger.
WebSocket clients can connect to the same path (plus /v1/messages and /v1/responses) with the same Authorization header, send the request body as one JSON text frame, and receive the streamed events as text frames until the socket closes.
- modelstringrequired
- A chat-capable id from /v1/models, e.g. claude-fable-5.1. Jev uses /v1/systemone instead.
- messagesarrayrequired
- Chat history: {role, content}. Content can be text or multimodal parts (vision models accept image parts).
- streamboolean
- true streams SSE chunks; the final chunk carries usage, including what the request cost.
- max_tokensinteger
- Output cap. Defaults to the model's maximum.
- temperaturenumber
- 0 to 2. Passed through to the model untouched.
- tools / tool_choicearray / string
- Standard OpenAI tool calling, passed through natively. No translation shims.
- response_formatobject
- e.g. {"type": "json_object"} on models that support JSON mode.
curl https://api.zenllm.org/v1/chat/completions \
-H "Authorization: Bearer $ZENLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-fable-5.1",
"messages": [{"role": "user", "content": "Hello"}],
"stream": true
}'Messages
Native Anthropic Messages format, including system prompts, thinking and tool use. Works unchanged with the official Anthropic SDKs. POST /v1/messages/count_tokens counts without running.
- modelstringrequired
- Any chat model id from the catalog.
- max_tokensintegerrequired
- Required by the Messages format.
- messagesarrayrequired
- User and assistant turns, Anthropic content blocks supported.
- systemstring
- System prompt, kept out of the messages array.
- thinkingobject
- Extended thinking config for models that support it.
- streamboolean
- Server-sent events in Anthropic's event framing.
curl https://api.zenllm.org/v1/messages \
-H "x-api-key: $ZENLLM_API_KEY" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-fable-5.1",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "Hello"}]
}'Responses
The OpenAI Responses format. The router converts between chat and responses dialects when a provider only supports one. Decision models use their native endpoint.
- modelstringrequired
- Model id. The router converts dialects when the upstream only speaks chat.
- inputstring | arrayrequired
- Prompt text or structured input items.
- streamboolean
- Streams response events as SSE.
- reasoningobject
- Effort hints for reasoning models that accept them.
curl https://api.zenllm.org/v1/responses \
-H "Authorization: Bearer $ZENLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-astra",
"input": "Summarize the attached notes.",
"stream": true
}'System One decisions
The TypeSafe System One API shape: send one state and a map of questions to Jev 1.13, then receive native typed answers with probabilities and confidence. This endpoint is non-streaming and does not accept chat messages, text-generation parameters, or image input.
The 64k-token total limit includes every question; the state plus the longest question must fit within 32k tokens. Input costs $0.021 / 1M tokens, 50% off TypeSafe's list price. Output tokens are free. The response keeps TypeSafe's native model, answers, and usage fields; check your ZenLLM ledger for the settled charge. Jev cannot be called through /v1/chat/completions or /v1/responses. See the TypeSafe API reference for full question semantics.
- modelstringrequired
- Use jev-1.13. Jev is a decision model, not a chat or text-generation model.
- statestring | object | arrayrequired
- Text or structured content to evaluate. One state is shared by all questions.
- questionsmap<string, question>required
- Named questions, each with a type and instructions. Answers use the same names. Multiple questions run in parallel.
- questions.*.typenoul | choice | scorerequired
- Noul returns a 0–1 yes probability. Choice selects an option with probabilities and confidence. Score returns a weighted level with probabilities and confidence.
- questions.*.instructionsstring | object | arrayrequired
- A specific, focused judgment. Structured values can refer to fields in state.
- questions.*.criteriaobject | array
- Choice: 1–255 named options. Score: 2–10 ordered levels. Noul: optional true/false descriptions.
curl https://api.zenllm.org/v1/systemone \
-H "Authorization: Bearer $ZENLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-1.13",
"state": {"ticket": "My card was charged twice; please help today."},
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {"billing": "Payments and refunds", "technical": "Bugs and outages"}
},
"is_urgent": {
"type": "noul",
"instructions": "Does the ticket need attention today?"
},
"severity": {
"type": "score",
"instructions": "How serious is the customer impact?",
"criteria": ["Low", "Moderate", "High"]
}
}
}'Embeddings and rerank
OpenAI-compatible embeddings plus POST /v1/rerank for scoring a document list against a query in retrieval pipelines.
- modelstringrequired
- text-embedding-3-large.
- inputstring | arrayrequired
- One string or a batch. Billed on total input tokens.
- top_ninteger
- Rerank only: how many documents to return, best first.
curl https://api.zenllm.org/v1/embeddings \
-H "Authorization: Bearer $ZENLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "text-embedding-3-large",
"input": "The quick brown fox"
}'Images
Image generation and edits, both JSON. Billing is exactly the listed rate, never more: generations cost n × the per-image price (FLUX: megapixels of output × the per-MP price), edits cost the flat edit price for the one image they return, and every response carries the settled usage.cost. If capacity is unavailable at these rates you get a retryable error, never a surprise charge. POST /v1/images/edits takes the source image as a URL or data: URI.
- modelstringrequired
- e.g. nano-banana-2, nano-banana-pro, flux.2-flex.
- promptstringrequired
- What to draw. For edits, how to change the input image.
- ninteger
- Number of images (1-10, default 1).
- sizestring
- e.g. 1024x1024. Model default when omitted.
- imagestring
- Edits only: source image as an https URL or data: URI.
- input_imagesarray
- Edits only: up to 8 source images; first is the base, rest are references.
curl https://api.zenllm.org/v1/images/generations \
-H "Authorization: Bearer $ZENLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "nano-banana-2",
"prompt": "isometric server rack, warm charcoal, orange accents",
"size": "1024x1024"
}'Audio
Text to speech returns raw audio bytes; POST /v1/audio/transcriptions takes JSON audio input and returns text.
- inputstringrequired
- Speech: the text to speak.
- voicestring
- Speech: voice preset, e.g. alloy.
- filestringrequired
- Transcriptions: audio as an https URL or data: URI (JSON body).
- languagestring
- Transcriptions: ISO hint, auto-detected when omitted.
curl https://api.zenllm.org/v1/audio/speech \
-H "Authorization: Bearer $ZENLLM_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "tts", "input": "Deploy is live.", "voice": "alloy"}' \
--output speech.mp3Models and health
The catalog with per-model pricing and context windows, no auth required. The detail route also lists the model's provider chain. GET /health reports service status.
GPT-6 Sol, GPT-5.6 Sol, and GPT-5.6 Terra request Azure Fast processing by default. Set service_tier: "default" to opt out. The nonstream response includes the tier actually served; streams send it in x-zenllm-service-tier. All three models are billed at their listed Standard token rates, whether Azure serves Fast or Standard. Fast does not apply to marketplace routes.
curl https://api.zenllm.org/v1/models
# per-model detail incl. provider chain
curl https://api.zenllm.org/v1/models/claude-fable-5.1
# service health
curl https://api.zenllm.org/health