Files
whit3rabbitandClaude Opus 4.8 e96ec18b24 feat: admin UI on by bare launch + tier-router review fixes
Admin/browser:
- Enable admin UI on bare (zero-arg) launch; auto-open default browser
  (main_helpers::bootstrap::admin_enabled / is_default_launch, browser.rs).
- Guard Docker (docker-entrypoint.sh) and systemd (packaging/anyllm-proxy.service)
  so headless server installs keep admin opt-in (DISABLE_ADMIN default off).

Claude Code tier router fixes (from code review):
- put.rs: validate only *enabled* tiers, and accept statically-configured
  (all_backends) targets via new SharedState.static_backends, not just managed.
- openai_signals: drop historical reasoning_content check so a plain follow-up
  in a reasoning conversation isn't misrouted to the Think tier.
- resolve_router_tier: warn! on fail-open when an active tier's backend is
  unknown instead of silently bypassing the router.

Tests: static-config backend acceptance; existing router coverage still green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-15 20:38:25 -05:00

605 lines
33 KiB
Markdown

# Environment Variables
## Env Files
Instead of setting variables in the shell, you can store them in a `.env` file and load it at startup.
**Auto-load:** If `.anyllm.env` exists in the current directory, it is loaded automatically. If not found, `~/.anyllm/.anyllm.env` is checked.
**Explicit flag:**
```bash
anyllm_proxy --env-file ~/configs/deepseek.env
```
**File format** (`KEY=VALUE`, Docker `--env-file` compatible):
```env
# Comments are supported
OPENAI_API_KEY=sk-...
OPENAI_BASE_URL=https://api.deepseek.com/v1
BIG_MODEL=deepseek-coder
SMALL_MODEL=deepseek-chat
export LISTEN_PORT=3000 # export prefix is also accepted
```
Rules:
- Lines starting with `#` are ignored.
- Values may be optionally quoted with `"double"` or `'single'` quotes.
- Double-quoted values interpret backslash escapes (`\n`, `\t`, `\r`, `\\`, `\"`).
- Single-quoted values are literal (no escape processing, matching bash behavior).
- Environment variables already set in the shell take precedence over the file.
- Variables previously imported via the admin UI (stored in SQLite) are applied after env files, with env files taking precedence.
- Use `docker run --env-file <path>` to pass the same file to a container.
The admin UI (Settings tab) has an **Export .env** button that generates a template from the current running configuration.
---
## Core
These are the variables most users need.
| Variable | Default | Description |
|----------|---------|-------------|
| `OPENAI_API_KEY` | (empty) | OpenAI API key. Required for the default `openai` backend. |
| `OPENAI_BASE_URL` | `https://api.openai.com` | Base URL for the upstream API. Change this to point at compatible APIs (Ollama, OpenRouter, etc.). Validated at startup (rejects private IPs, loopback, cloud metadata endpoints). |
| `OPENAI_API_FORMAT` | `chat` | Which OpenAI API format to use. `chat` (default) for Chat Completions, `responses` for the Responses API. Only relevant when `BACKEND=openai`. |
| `BACKEND` | `openai` | Which upstream backend to target. Valid values: `openai`, `azure`, `vertex`, `gemini`, `anthropic`, `bedrock`. |
| `LISTEN_PORT` | `3000` | Port the proxy listens on. |
| `BIG_MODEL` | (per backend) | Model used when the request specifies a sonnet or opus model. Defaults: `gpt-4o` (openai/azure), `gemini-2.5-pro` (vertex/gemini), Bedrock model ID (bedrock). Not used for `anthropic` backend (passthrough). |
| `SMALL_MODEL` | (per backend) | Model used when the request specifies a haiku model. Defaults: `gpt-4o-mini` (openai/azure), `gemini-2.5-flash` (vertex/gemini), Bedrock model ID (bedrock). |
| `RUST_LOG` | `info` | Tracing filter. Examples: `debug`, `anyllm_proxy=trace`. |
| `LOG_BODIES` | `false` | Log request/response bodies at debug level. Set to `true` or `1`. **Warning:** may expose sensitive data (prompts, API keys, PII). |
| `REDACT_SECRETS` | `false` | Scan upstream JSON/text request payloads and replace detected secrets before forwarding. Set to `true` or `1`. Also available as `--redact-secrets` and in the admin UI. |
| `ANYLLM_DEGRADATION_WARNINGS` | `false` | Expose `x-anyllm-degradation` response header when features are silently dropped during translation. Set to `true` or `1`. Automatically enabled when `PROXY_CONFIG` is set. |
| `DISABLE_ADMIN` | (unset) | Set to `1`, `true`, or `yes` to force-disable the admin web interface even when `--webui` is passed. Useful in automated/container environments. |
## Auth
| Variable | Default | Description |
|----------|---------|-------------|
| `PROXY_API_KEYS` | (unset) | Comma-separated list of allowed API keys. Clients must send one of these as their Bearer token. If unset and `PROXY_OPEN_RELAY` is not set, all requests are rejected with 401. |
| `PROXY_OPEN_RELAY` | (unset) | Set to `true` or `1` to accept any non-empty API key. **Local dev only.** Logged as an error when bound to a non-loopback address. |
| `PROXY_CONFIG` | (unset) | Path to a config file (simple YAML, LiteLLM YAML, or TOML). Auto-detected from `~/.anyllm/config.yaml` if not set. See [CONFIG.md](CONFIG.md). |
## Network / Security
| Variable | Default | Description |
|----------|---------|-------------|
| `IP_ALLOWLIST` | (unset) | Comma-separated list of allowed client IPs or CIDR ranges (e.g. `10.0.0.0/8,192.168.1.5`). When set, requests from other IPs are rejected. |
| `TRUST_PROXY_HEADERS` | `false` | Trust `X-Forwarded-For` and `X-Real-IP` headers for client IP resolution. Set to `true` or `1` when behind a reverse proxy. |
| `REQUEST_TIMEOUT_SECS` | `900` | Wall-clock cap (seconds) for streaming responses. 0 = disabled. |
| `OMIT_STREAM_OPTIONS` | `false` | Strip `stream_options` from streaming requests. Needed for local LLMs (older Ollama, text-generation-webui, LM Studio) that reject unknown fields with HTTP 400. |
## Tool Guardrails
These apply only when a config file initializes the tool engine through `tool_execution`, `builtin_tools`, or `mcp_servers`.
| Variable | Default | Description |
|----------|---------|-------------|
| `FORGE_TOOL_CALL_POLICY` | `disabled` | Set to `standard` to enable Forge-style advisory guardrails for model-produced tool calls. YAML `tool_execution.guardrails` takes precedence. |
## Prompt Compression (Optimizer, optional)
Opt-in Frozen-Frontier Extractive Compression (FFEC) of long client-sent conversation
history, applied at the parsed-body seam (OpenAI Chat Completions and Anthropic Messages)
before translation. Never touches proxy-internal tool-loop turns, only what the client sent.
| Variable | Default | Description |
|----------|---------|-------------|
| `OPTIMIZER_MODE` | `off` | `off`, `shadow`, or `live`. `shadow` runs the full pipeline and logs an `OptimizationReport` (would-be token savings) but forwards the original body unchanged. `live` renders the compressed body back in place and, for Anthropic requests, places a `cache_control` breakpoint at the compression frontier. Seeds `RuntimeConfig.optimizer_mode` (also live-toggleable from the admin UI / `PUT /admin/api/config` without a restart). |
Fails open on any error (malformed body, panic in the adapter/algorithm pipeline): the
original request is forwarded unchanged and the mode is reported as a no-op. See
`crates/optimizer/CLAUDE.md` for the compression algorithm and `record_optimization`
counters (`optimizer_compressed_total`, `optimizer_messages_compressed_total`,
`optimizer_removed_tokens_total`) exposed on `GET /metrics`.
### ONNX scorer (opt-in, LLMLingua-2)
Live mode uses a heuristic scorer by default. Build the proxy with `--features
optimizer-onnx` to enable the LLMLingua-2 ONNX scorer (`ort` downloads an onnxruntime
binary at build time). The ~170MB model is never bundled or auto-downloaded: fetch it once
from the admin UI (Settings → Prompt compression → **Download model**, which
sha256-verifies against a pinned digest) or with the `optimize-model` CLI. The verified
`model.onnx` + `tokenizer.json` land in `<ANYLLM_HOME>/models/<sha256>/`; the proxy loads
the scorer eagerly at startup if present, else lazily on the first live request after a
download. Admin endpoints: `GET /admin/api/optimizer/model` (status), `POST` (start
download).
| Variable | Default | Description |
|----------|---------|-------------|
| `MODEL_URL` | pinned HF repo | Base URL the artifact (`<url>/model.onnx`, `<url>/tokenizer.json`) is fetched from. |
| `MODEL_SHA256` | pinned digest | sha256 the downloaded `model.onnx` must match; the download is rejected on mismatch. |
| `MODEL_CACHE_DIR` | `<ANYLLM_HOME>/models` | Cache root; the verified pair lands in `<dir>/<sha256>/`. |
## OIDC / JWT Authentication (optional)
When `OIDC_ISSUER_URL` is set, the proxy discovers the OIDC configuration and loads JWKS. Tokens that look like JWTs are validated against the JWKS before falling through to key-based auth.
| Variable | Default | Description |
|----------|---------|-------------|
| `OIDC_ISSUER_URL` | (unset) | OIDC issuer URL for JWT validation (e.g. `https://accounts.google.com`). Enables OIDC authentication when set. |
| `OIDC_AUDIENCE` | (issuer URL) | Expected audience claim in JWTs. Defaults to the issuer URL if not set. |
## AWS Bedrock
Set `BACKEND=bedrock` to route through AWS Bedrock. The proxy sends Anthropic Messages API format directly to Bedrock (no OpenAI translation). Requests are signed with AWS SigV4.
| Variable | Default | Description |
|----------|---------|-------------|
| `AWS_REGION` | (required) | AWS region, e.g. `us-east-1`. |
| `AWS_ACCESS_KEY_ID` | (required) | AWS access key ID for SigV4 signing. |
| `AWS_SECRET_ACCESS_KEY` | (required) | AWS secret access key for SigV4 signing. |
| `AWS_SESSION_TOKEN` | (optional) | Temporary session token for STS credentials. |
| `BIG_MODEL` | `anthropic.claude-sonnet-4-20250514-v1:0` | Bedrock model ID for sonnet/opus requests. |
| `SMALL_MODEL` | `anthropic.claude-haiku-4-5-20251001-v1:0` | Bedrock model ID for haiku requests. |
### Example
```bash
BACKEND=bedrock \
AWS_REGION=us-east-1 \
AWS_ACCESS_KEY_ID=AKIA... \
AWS_SECRET_ACCESS_KEY=wJalr... \
cargo run -p anyllm_proxy
```
### Streaming
Bedrock streaming uses AWS Event Stream binary framing instead of SSE. The proxy decodes Event Stream frames and re-emits them as standard SSE events, so downstream clients see the same Anthropic SSE format as with other backends.
---
## Azure OpenAI
Set `BACKEND=azure` to route through Azure OpenAI Service. The request/response format is identical to standard OpenAI Chat Completions; only the URL scheme and auth header differ.
| Variable | Default | Description |
|----------|---------|-------------|
| `AZURE_OPENAI_API_KEY` | (required) | Azure OpenAI API key. Sent as `api-key` header. |
| `AZURE_OPENAI_ENDPOINT` | (required) | Full Azure resource endpoint, e.g. `https://my-resource.openai.azure.com`. Accepts sovereign cloud URLs. |
| `AZURE_OPENAI_DEPLOYMENT` | (required) | Deployment name (the model deployment you created in Azure portal). |
| `AZURE_OPENAI_API_VERSION` | `2024-10-21` | Azure API version string appended as `?api-version=` query parameter. |
The proxy constructs the full URL as:
```
{AZURE_OPENAI_ENDPOINT}/openai/deployments/{AZURE_OPENAI_DEPLOYMENT}/chat/completions?api-version={AZURE_OPENAI_API_VERSION}
```
### Example
```bash
BACKEND=azure \
AZURE_OPENAI_API_KEY=abc123 \
AZURE_OPENAI_ENDPOINT=https://my-resource.openai.azure.com \
AZURE_OPENAI_DEPLOYMENT=gpt-4o \
cargo run -p anyllm_proxy
```
---
## Google Vertex AI
Set `BACKEND=vertex` to route through Google Vertex AI. The proxy constructs the Vertex AI endpoint URL from the project and region, then forwards via the OpenAI-compatible API.
| Variable | Default | Description |
|----------|---------|-------------|
| `VERTEX_PROJECT` | (required) | GCP project ID. |
| `VERTEX_REGION` | (required) | GCP region, e.g. `us-central1`. |
| `VERTEX_API_KEY` | (one required) | Google API key for authentication. Either this or `GOOGLE_ACCESS_TOKEN` must be set. |
| `GOOGLE_ACCESS_TOKEN` | (one required) | OAuth2 access token for authentication. Alternative to `VERTEX_API_KEY`. |
| `BIG_MODEL` | `gemini-2.5-pro` | Model for sonnet/opus requests. |
| `SMALL_MODEL` | `gemini-2.5-flash` | Model for haiku requests. |
The proxy constructs the endpoint as:
```
https://{VERTEX_REGION}-aiplatform.googleapis.com/v1/projects/{VERTEX_PROJECT}/locations/{VERTEX_REGION}/endpoints/openapi
```
### Example
```bash
BACKEND=vertex \
VERTEX_PROJECT=my-project \
VERTEX_REGION=us-central1 \
VERTEX_API_KEY=AIza... \
cargo run -p anyllm_proxy
```
---
## Google Gemini
Set `BACKEND=gemini` to route through the Gemini API (generativelanguage.googleapis.com). Uses the OpenAI-compatible endpoint.
| Variable | Default | Description |
|----------|---------|-------------|
| `GEMINI_API_KEY` | (required) | Gemini API key. Sent as `x-goog-api-key` header. |
| `GEMINI_BASE_URL` | `https://generativelanguage.googleapis.com/v1beta` | Base URL. The proxy appends `/openai` to reach the OpenAI-compatible endpoint. |
| `BIG_MODEL` | `gemini-2.5-pro` | Model for sonnet/opus requests. |
| `SMALL_MODEL` | `gemini-2.5-flash` | Model for haiku requests. |
### Example
```bash
BACKEND=gemini \
GEMINI_API_KEY=AIza... \
cargo run -p anyllm_proxy
```
---
## Anthropic Passthrough
Set `BACKEND=anthropic` to forward Anthropic Messages API requests directly to the Anthropic API without any translation. Model names are passed through unchanged (no BIG_MODEL/SMALL_MODEL mapping).
| Variable | Default | Description |
|----------|---------|-------------|
| `ANTHROPIC_API_KEY` | (required) | Anthropic API key. |
| `ANTHROPIC_BASE_URL` | `https://api.anthropic.com` | Base URL for the Anthropic API. |
| `ANTHROPIC_THINKING_REPAIR` | `false` | Repair corrupted `thinking`/`redacted_thinking` blocks in the last assistant message of `/v1/messages` requests before forwarding upstream. See below. |
| `ANTHROPIC_FORWARD_CLIENT_AUTH` | `false` | Forward the client's own `x-api-key`/`Authorization` header upstream verbatim instead of `ANTHROPIC_API_KEY`/`ANTHROPIC_AUTH_TOKEN`. See below. |
| `PXPIPE_COMPRESS` | `false` | Enable text-to-image context compression on `/v1/messages` (`BACKEND=anthropic` passthrough): render the stable system+tools slab to a PNG and swap it in to save input tokens on vision models. Also settable via `pxpipe_compress: true` in simple YAML config, and live-toggleable from the admin config API. See below. |
| `PXPIPE_HISTORY` | `false` | Additionally collapse the OLD closed-tool-call conversation prefix into history image(s) (keeping the recent tail as text). Off by default: highest cache-stability risk of the feature. Only meaningful when `PXPIPE_COMPRESS=true`. |
| `PXPIPE_MODELS` | `claude-fable-5` | CSV of model bases in scope for `PXPIPE_COMPRESS` (substring match). **Seeds** the default scope; the runtime scope is then editable per-model from the admin UI (Settings tab shows the vision-capable models as checkboxes). Out-of-scope or non-vision models pass through untouched. Default is conservative because weaker readers (e.g. Opus 4.8) degrade on imaged content. |
### Example
```bash
BACKEND=anthropic \
ANTHROPIC_API_KEY=sk-ant-... \
cargo run -p anyllm_proxy
```
### Thinking-block repair (`ANTHROPIC_THINKING_REPAIR=true`)
Clients that replay conversation history (e.g. Claude Code) can corrupt the
`thinking`/`redacted_thinking` blocks in the last assistant message — merged
text from interleaved streams, dropped `redacted_thinking` blocks that never
get persisted to disk, reordered blocks. The Anthropic API validates those
blocks byte-exactly against their signatures, so any mutation produces a
repeating 400 until the client's context is cleared.
With this flag set, the proxy records every response's content blocks
(text, signatures, `redacted_thinking` data, `tool_use` ownership) as ground
truth, then on each outgoing request verifies and repairs only the *last*
assistant message against it: byte-identical blocks pass through untouched,
blocks with a known signature but mutated text are restored to the recorded
original, and blocks belonging to a different recorded message ("intruder"
blocks) are dropped. Messages before the last assistant one are never
touched, so prompt-cache prefixes are preserved.
The ground-truth store is in-memory only (bounded, no persistence). On
proxy restart it starts empty; requests are forwarded unrepaired until a
fresh response is recorded, then repair resumes on the next turn. Off by
default; only takes effect for `BACKEND=anthropic` passthrough (`/v1/messages`).
This can also be toggled live from the admin UI (Settings tab) or via
`PUT /admin/api/config` with `{"anthropic_thinking_repair": true|false}` — no
restart required. The env var only sets the value at startup; an admin-UI
change takes effect immediately and persists across restarts via SQLite until
reset.
### Forwarding the client's own credential (`ANTHROPIC_FORWARD_CLIENT_AUTH=true`)
By default the proxy always sends **its own** `ANTHROPIC_API_KEY`/
`ANTHROPIC_AUTH_TOKEN` to the real Anthropic API — the client's incoming
`x-api-key`/`Authorization` header is only ever checked against the proxy's
own inbound auth (`PROXY_API_KEYS`/`PROXY_OPEN_RELAY`) and then discarded.
Setting this flag instead forwards that exact header — same name, same
value, byte-for-byte, no re-shaping — upstream in place of the operator's
configured credential. This lets Claude Code use its own Pro/Max
subscription OAuth session directly through the proxy, without a separate
`claude setup-token` step.
Since the credential that authenticates a request into the proxy becomes the
literal credential sent to Anthropic, this only makes sense for a
single-key/BYOK deployment where those two are meant to be the same thing.
It is automatically skipped (the operator's own credential is used instead,
regardless of the flag) for any request authenticated via a virtual key or
OIDC/JWT — a virtual key is deliberately not a real Anthropic credential, and
forwarding a JWT upstream would never work. A client that authenticated via
the Gemini-CLI-compatible `x-goog-api-key` header has its value forwarded
renamed to `x-api-key` (the only credential header name Anthropic itself
understands), not literally as `x-goog-api-key`.
At startup, the proxy refuses to start with this flag on if `PROXY_API_KEYS`
has 2+ distinct entries and `PROXY_OPEN_RELAY` is not set, since that
combination would let different callers each redirect the upstream Anthropic
credential. The same rule is enforced live: this flag is toggleable from the
admin UI (**Settings**) or `PUT /admin/api/config` with no restart, and that
route rejects the same misconfigured combination with a 400 rather than
silently accepting it.
```bash
BACKEND=anthropic \
ANTHROPIC_AUTH_TOKEN=$(claude setup-token) \
ANTHROPIC_FORWARD_CLIENT_AUTH=true \
PROXY_OPEN_RELAY=true \
anyllm_proxy
```
Only active for `BACKEND=anthropic` passthrough (`/v1/messages` and the
generic Anthropic-native catch-all route). Off by default. Applies uniformly
to every `BackendKind::Anthropic` backend in a multi-backend deployment (one
shared runtime setting, like `ANTHROPIC_THINKING_REPAIR`) rather than being
configurable per backend.
---
## Third-party OpenAI-compatible providers
Any LiteLLM provider id from the built-in catalog can be used as a `BACKEND` value. These providers use the OpenAI Chat Completions protocol and route through the same HTTP client as `BACKEND=openai`. Legacy local IDs such as `gmi_cloud`, `public_ai`, `zhipuai`, `ai_ml_api`, `github`, `jina`, `exa`, and `stability_ai` are accepted only as migration aliases.
**Resolution order:**
1. `BACKEND=<provider_id>` — e.g. `BACKEND=groq`
2. Base URL: `OPENAI_BASE_URL` env var (if set) overrides the provider default; otherwise the catalog default is used. Known providers without a safe global default require `OPENAI_BASE_URL`.
3. API key: `OPENAI_API_KEY` (if set) takes precedence; otherwise the provider-specific key var is used (e.g. `GROQ_API_KEY`).
**Example (Groq):**
```bash
BACKEND=groq \
GROQ_API_KEY=gsk_... \
PROXY_OPEN_RELAY=true \
cargo run -p anyllm_proxy
```
### Popular cloud providers
| `BACKEND` value | API key env var(s) | Default base URL |
|---|---|---|
| `groq` | `GROQ_API_KEY` | `https://api.groq.com/openai/v1` |
| `together_ai` | `TOGETHER_API_KEY`, `TOGETHERAI_API_KEY` | `https://api.together.xyz/v1` |
| `openrouter` | `OPENROUTER_API_KEY` | `https://openrouter.ai/api/v1` |
| `fireworks_ai` | `FIREWORKS_API_KEY` | `https://api.fireworks.ai/inference/v1` |
| `mistral` | `MISTRAL_API_KEY` | `https://api.mistral.ai/v1` |
| `codestral` | `CODESTRAL_API_KEY` | `https://codestral.mistral.ai/v1` |
| `perplexity` | `PERPLEXITYAI_API_KEY`, `PERPLEXITY_API_KEY` | `https://api.perplexity.ai` |
| `deepseek` | `DEEPSEEK_API_KEY` | `https://api.deepseek.com` |
| `cerebras` | `CEREBRAS_API_KEY` | `https://api.cerebras.ai/v1` |
| `xai` | `XAI_API_KEY` | `https://api.x.ai/v1` |
| `nvidia_nim` | `NVIDIA_NIM_API_KEY` | `https://integrate.api.nvidia.com/v1` |
| `sambanova` | `SAMBANOVA_API_KEY` | `https://api.sambanova.ai/v1` |
| `nebius` | `NEBIUS_API_KEY` | `https://api.studio.nebius.ai/v1` |
| `deepinfra` | `DEEPINFRA_API_KEY` | `https://api.deepinfra.com/v1/openai` |
| `novita` | `NOVITA_API_KEY` | `https://api.novita.ai/v3/openai` |
| `hyperbolic` | `HYPERBOLIC_API_KEY` | `https://api.hyperbolic.xyz/v1` |
| `lambda_ai` | `LAMBDA_API_KEY` | `https://api.lambdalabs.com/v1` |
| `nscale` | `NSCALE_API_KEY` | `https://inference.nscale.com/v1` |
| `featherless_ai` | `FEATHERLESS_API_KEY` | `https://api.featherless.ai/v1` |
| `friendliai` | `FRIENDLIAI_TOKEN` | `https://api.friendli.ai/serverless/v1` |
| `replicate` | `REPLICATE_API_KEY` | `https://openai-compat.replicate.com/v1` |
| `cohere_chat` | `COHERE_API_KEY` | `https://api.cohere.com/compatibility/v1` |
| `ai21` | `AI21_API_KEY` | `https://api.ai21.com/studio/v1` |
| `anyscale` | `ANYSCALE_API_KEY` | `https://api.endpoints.anyscale.com/v1` |
| `aleph_alpha` | `ALEPH_ALPHA_API_KEY` | `https://api.aleph-alpha.com` |
| `nlp_cloud` | `NLP_CLOUD_API_KEY` | `https://api.nlpcloud.io` |
| `clarifai` | `CLARIFAI_API_KEY` | `https://api.clarifai.com/v2` |
| `predibase` | `PREDIBASE_API_KEY` | `https://serving.app.predibase.com` |
| `voyage` | `VOYAGE_API_KEY` | `https://api.voyageai.com/v1` (embeddings only) |
| `jina_ai` | `JINA_AI_API_KEY` | `https://api.jina.ai/v1` (embeddings/rerank) |
| `github_copilot` | `GITHUB_TOKEN` | `https://models.github.ai/inference` |
| `chutes` | `CHUTES_API_KEY` | `https://llm.chutes.ai/v1` |
| `gmi` | `GMI_CLOUD_API_KEY` | `https://api.gmi-serving.com/v1` |
| `meta_llama` | `META_LLAMA_API_KEY` | `https://www.llama.com/api/v1` |
| `aiml` | `AIML_API_KEY` | `https://api.aimlapi.com/v1` |
| `morph` | `MORPH_API_KEY` | `https://api.morphllm.com/v1` |
| `galadriel` | `GALADRIEL_API_KEY` | `https://api.galadriel.com/v1` |
| `nanogpt` | `NANOGPT_API_KEY` | `https://nano-gpt.com/api/v1` |
| `bytez` | `BYTEZ_KEY` | `https://api.bytez.com/models/v2` |
| `publicai` | `PUBLIC_AI_API_KEY` | `https://api.publicai.co/v1` |
### Regional / specialized
| `BACKEND` value | API key env var(s) | Default base URL |
|---|---|---|
| `moonshot` | `MOONSHOT_API_KEY` | `https://api.moonshot.cn/v1` |
| `volcengine` | `VOLCENGINE_API_KEY` | `https://ark.cn-beijing.volces.com/api/v3` |
| `minimax` | `MINIMAX_API_KEY` | `https://api.minimax.chat/v1` |
| `zai` | `ZHIPUAI_API_KEY` | `https://open.bigmodel.cn/api/paas/v4` |
| `dashscope` | `DASHSCOPE_API_KEY` | `https://dashscope.aliyuncs.com/compatible-mode/v1` |
| `xiaomi_mimo` | `XIAOMI_MIMO_API_KEY` | `https://api.mimo.chat/v1` |
| `gradient_ai` | `GRADIENT_ACCESS_TOKEN` | `https://api.gradient.ai` |
### Per-deployment (must set `OPENAI_BASE_URL`)
These providers require a workspace/account-specific URL set via `OPENAI_BASE_URL`.
| `BACKEND` value | API key env var(s) | Notes |
|---|---|---|
| `databricks` | `DATABRICKS_API_KEY` | Set `OPENAI_BASE_URL` to your workspace serving endpoint |
| `hosted_vllm` | `VLLM_API_KEY` | Set `OPENAI_BASE_URL` to your vLLM server URL |
| `huggingface` | `HUGGINGFACE_API_KEY`, `HF_TOKEN` | Set `OPENAI_BASE_URL` to your HF Inference Endpoint |
| `scaleway` | `SCW_SECRET_KEY` | Set `OPENAI_BASE_URL` to your Scaleway Inference endpoint |
| `baseten` | `BASETEN_API_KEY` | Set `OPENAI_BASE_URL` to your Baseten deployment URL |
| `azure_ai` | `AZURE_AI_API_KEY`, `AZURE_AI_API_BASE` | Azure AI Foundry; also set `OPENAI_BASE_URL=<AZURE_AI_API_BASE>` |
| `watsonx` | `WATSONX_API_KEY`, `WATSONX_URL` | IBM WatsonX; also set `OPENAI_BASE_URL=<WATSONX_URL>` |
| `cloudflare` | `CLOUDFLARE_API_KEY`, `CLOUDFLARE_ACCOUNT_ID` | Set `OPENAI_BASE_URL` to your account endpoint |
| `snowflake` | `SNOWFLAKE_JWT`, `SNOWFLAKE_ACCOUNT_ID` | Set `OPENAI_BASE_URL` to your Snowflake Cortex endpoint |
| `xinference` | `XINFERENCE_SERVER_URL` | Set `OPENAI_BASE_URL=<XINFERENCE_SERVER_URL>` |
| `ovhcloud` | `OVH_AI_ENDPOINTS_ACCESS_TOKEN` | Set `OPENAI_BASE_URL` to your OVHcloud endpoint |
| `wandb` | `WANDB_API_KEY` | Set `OPENAI_BASE_URL` to your W&B Inference project URL |
### Self-hosted / local (no key required)
| `BACKEND` value | Default base URL | Notes |
|---|---|---|
| `ollama` | `http://localhost:11434/v1` | Override with `OPENAI_BASE_URL` for remote Ollama |
| `lm_studio` | `http://localhost:1234/v1` | LM Studio local server |
| `llamafile` | `http://localhost:8080` | llamafile server |
| `lemonade` | `http://localhost:8000` | Lemonade local server |
| `docker_model_runner` | `http://localhost:12434/engines/llama.cpp/v1` | Docker Model Runner |
| `infinity` | `http://localhost:7997` | Infinity embeddings server |
| `petals` | `http://localhost:8080` | Petals distributed inference |
| `triton` | (none — set `OPENAI_BASE_URL`) | NVIDIA Triton Inference Server |
> **Note:** `sagemaker` appears in the provider catalog but uses a custom AWS signing protocol not yet routed through `BACKEND=sagemaker`. Use `BACKEND=bedrock` for AWS-hosted Anthropic models instead.
---
## mTLS Client Certificates
Most users do not need these. They configure mutual TLS (mTLS) on the **outbound** connection from the proxy to the backend endpoint. Use them when the backend requires a client certificate for authentication, or uses a private CA that is not in the system trust store.
These variables do not affect the proxy's own listener. The proxy always serves plain HTTP. For inbound TLS termination, place a reverse proxy (nginx, caddy, etc.) in front.
| Variable | Default | Description |
|----------|---------|-------------|
| `TLS_CLIENT_CERT_P12` | (unset) | Path to a PKCS#12 (.p12 or .pfx) client certificate file. When set, the proxy presents this certificate during the TLS handshake with the backend. |
| `TLS_CLIENT_CERT_PASSWORD` | (unset) | Password to decrypt the P12 file. **Required** if `TLS_CLIENT_CERT_P12` is set. The proxy will refuse to start if the P12 is set without a password. |
| `TLS_CA_CERT` | (unset) | Path to a PEM-encoded CA certificate. Added to the trust store for verifying the backend's server certificate. Use this when the backend uses a private or self-signed CA. |
All three are optional. When unset, the proxy connects using the system's default TLS configuration and trust store.
### Validation
All certificate files are read and validated at startup. The proxy will panic with a descriptive error if:
- The P12 file does not exist or cannot be read.
- The P12 password is wrong or the file is corrupt.
- The CA certificate file does not exist or is not valid PEM.
- `TLS_CLIENT_CERT_P12` is set without `TLS_CLIENT_CERT_PASSWORD`.
### Example
```bash
OPENAI_API_KEY=sk-... \
OPENAI_BASE_URL=https://internal-llm.corp.example.com \
TLS_CLIENT_CERT_P12=/etc/proxy/client.p12 \
TLS_CLIENT_CERT_PASSWORD=changeit \
TLS_CA_CERT=/etc/proxy/corp-ca.pem \
cargo run -p anyllm_proxy
```
---
## Admin Web UI
The admin web interface starts **by default when the proxy is run with no arguments** (`anyllm_proxy`), and the default browser is opened to the admin page automatically. Passing any argument other than `--webui`/`--admin` (e.g. `--env-file`, `--redact-secrets`) keeps the proxy CLI-only. To run the admin UI alongside other flags without auto-opening a browser, pass `--webui` or `--admin` explicitly. The `WEBUI=1` or `ADMIN=1` environment variables also force it on (used by docker-entrypoint.sh). Set `DISABLE_ADMIN=1` to force it off in all cases.
```bash
anyllm_proxy --webui
```
The dashboard binds to `127.0.0.1:3001` by default (not externally accessible). It shows live request logs, latency percentiles, error rates, per-backend metrics, and lets you change log level and model mappings without restarting the server. The Settings tab also displays all active environment variables (secrets are masked).
| Variable | Default | Description |
|----------|---------|-------------|
| `ADMIN_PORT` | `3001` | Port for the admin dashboard. Must differ from `LISTEN_PORT`. |
| `ADMIN_BIND` | `127.0.0.1` | Bind address for the admin dashboard. Set to `0.0.0.0` to make it reachable from outside the host (required in Docker). |
| `ADMIN_TOKEN` | (generated) | Bearer token for the admin API. If unset, a random 256-bit hex token is generated at startup and written to `ADMIN_TOKEN_PATH`. On non-Unix platforms, auto-generation is not supported; set this explicitly. |
| `ADMIN_TOKEN_PATH` | `~/.anyllm/.admin_token` | File path where the generated admin token is written. Permissions are set to `0600` on Unix. |
| `ADMIN_DB_PATH` | `~/.anyllm/admin.db` | SQLite database path for request logging, config overrides, virtual keys, and model deployments. Config overrides survive restarts. |
| `ADMIN_LOG_RETENTION_DAYS` | `7` | Days to retain request log entries before automatic purge. |
| `DISABLE_ADMIN` | (unset) | Set to `1`, `true`, or `yes` to force-disable the admin server even when `--webui` is passed. Useful in container deployments where the flag might be baked into the entrypoint. |
### Token security
The admin token is written to `ADMIN_TOKEN_PATH` (default `~/.anyllm/.admin_token`) rather than stderr, because container log drivers capture stderr and persist it in centralized logging systems. On Unix, the file is created with mode `0600`. The token is printed to stdout for easy copy on first launch.
In production, set `ADMIN_TOKEN` explicitly:
```bash
ADMIN_TOKEN=$(openssl rand -hex 32) anyllm_proxy --webui
```
### Example
```bash
# Proxy + admin UI on a custom port with a fixed token
ADMIN_PORT=4000 \
ADMIN_TOKEN=my-secret-token \
ADMIN_DB_PATH=/var/lib/anyllm/admin.db \
anyllm_proxy --webui
# Open: http://127.0.0.1:4000/admin/?token=my-secret-token
```
---
## Webhooks / Callbacks
| Variable | Default | Description |
|----------|---------|-------------|
| `WEBHOOK_URLS` | (unset) | Comma-separated list of webhook URLs to POST request completion events to. |
| `BATCH_WEBHOOK_URLS` | (unset) | Comma-separated global webhook URLs for batch API job completions. Only active when admin is enabled. |
| `BATCH_WEBHOOK_SIGNING_SECRET` | (unset) | Secret for HMAC-signing batch webhook payloads. |
---
## Langfuse Integration (optional)
Send LLM generation events to Langfuse's batch ingestion API. Activated when both `LANGFUSE_PUBLIC_KEY` and `LANGFUSE_SECRET_KEY` are set, or when `"langfuse"` appears in `litellm_settings.callbacks` in a LiteLLM config file.
| Variable | Default | Description |
|----------|---------|-------------|
| `LANGFUSE_PUBLIC_KEY` | (required) | Langfuse public key. |
| `LANGFUSE_SECRET_KEY` | (required) | Langfuse secret key. |
| `LANGFUSE_HOST` | `https://cloud.langfuse.com` | Langfuse API host. Validated against SSRF (rejects private IPs). |
---
## Distributed Rate Limiting (optional)
Requires building with `--features redis`. When `REDIS_URL` is set, RPM/TPM rate limit checks are performed against Redis so multiple proxy instances share rate limit state.
| Variable | Default | Description |
|----------|---------|-------------|
| `REDIS_URL` | (unset) | Redis connection URL (e.g. `redis://localhost:6379`). Enables distributed rate limiting when set. |
| `RATE_LIMIT_FAIL_POLICY` | `open` | Behavior when Redis is unreachable: `open` (allow requests through) or `closed`/`deny` (reject requests). |
---
## Semantic Cache (optional)
Requires building with `--features qdrant`. Uses Qdrant for embedding-based response caching.
| Variable | Default | Description |
|----------|---------|-------------|
| `QDRANT_URL` | (unset) | Qdrant connection URL. Enables semantic caching when set. |
| `QDRANT_COLLECTION` | (unset) | Qdrant collection name for cached responses. |
---
## Cost Tracking
| Variable | Default | Description |
|----------|---------|-------------|
| `MODEL_PRICING_FILE` | (embedded) | Path to a JSON file overriding the embedded model pricing data. |
---
## OpenTelemetry (optional)
Trace export is opt-in. Build with the `otel` cargo feature to enable it:
```bash
cargo build -p anyllm_proxy --features otel
```
When the feature is enabled, the proxy initializes an OTLP span exporter that sends traces over HTTP/protobuf. The OTLP SDK reads configuration from standard environment variables; no proxy-specific config is needed.
| Variable | Default | Description |
|----------|---------|-------------|
| `OTEL_EXPORTER_OTLP_ENDPOINT` | `http://localhost:4318` | OTLP collector endpoint (HTTP). |
| `OTEL_SERVICE_NAME` | `unknown_service` | Service name attached to all exported spans. Set this to `anyllm-proxy` or your deployment name. |
| `OTEL_TRACES_SAMPLER` | `parentbased_always_on` | Sampling strategy. Common values: `always_on`, `always_off`, `traceidratio` (pair with `OTEL_TRACES_SAMPLER_ARG`). |
| `OTEL_TRACES_SAMPLER_ARG` | (none) | Argument for the sampler, e.g. `0.1` for 10% sampling with `traceidratio`. |
When built without the `otel` feature (the default), none of these variables have any effect and there is zero runtime overhead.
---
## LiteLLM Environment Variable Aliases
For compatibility with LiteLLM configurations, the proxy recognizes these aliases. Aliases only take effect when the target variable is not already set.
| LiteLLM Variable | Maps To |
|------------------|---------|
| `LITELLM_MASTER_KEY` | `PROXY_API_KEYS` |
| `LITELLM_CONFIG` | `PROXY_CONFIG` |
| `AZURE_API_KEY` | `AZURE_OPENAI_API_KEY` |
| `AZURE_API_BASE` | `AZURE_OPENAI_ENDPOINT` |
| `AZURE_API_VERSION` | `AZURE_OPENAI_API_VERSION` |
| `AWS_REGION_NAME` | `AWS_REGION` |
| `LITELLM_IP_ALLOWLIST` | `IP_ALLOWLIST` |