Document ANTHROPIC_AUTH_TOKEN, ANTHROPIC_API_KEY="", and CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 alongside ANTHROPIC_BASE_URL so users know the full set of vars needed to run Claude Code against the proxy. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
anyllm-proxy
An API translation proxy that lets Anthropic-based tools (Claude Code, Cursor, Windsurf, Cline) talk to any OpenAI-compatible backend, local LLM, or alternative provider.
Quick Start
Download a binary from the releases page, or install from source:
cargo install anyllm_proxy
Create a .anyllm.env config file:
OPENAI_API_KEY=unused
OPENAI_BASE_URL=http://localhost:11434/v1
BIG_MODEL=qwen2.5-coder:32b
SMALL_MODEL=qwen2.5-coder:32b
Run the proxy (auto-loads .anyllm.env from the current directory):
anyllm_proxy
# or: anyllm_proxy --env-file ~/configs/ollama.env
Simple mode vs. advanced mode
| Simple mode | Advanced mode | |
|---|---|---|
| Config | 3 env vars or .anyllm.env |
config.toml / config.yaml |
| Routing | Single backend | Multi-backend with path prefixes |
| Admin UI | Not started | --webui flag |
| Translation warnings | Silent (never exposed to clients) | x-anyllm-degradation header active |
| How to enable | Default | Pass --webui, set PROXY_CONFIG, or ANYLLM_DEGRADATION_WARNINGS=true |
Most users never leave simple mode. Start there.
Point Claude Code at the proxy:
ANTHROPIC_BASE_URL=http://localhost:3000 \
ANTHROPIC_AUTH_TOKEN=proxy-user \
ANTHROPIC_API_KEY="" \
CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 \
claude
Admin Web Interface (optional)
Pass --webui (or --admin) to also start the admin dashboard on 127.0.0.1:3001. The dashboard has the following tabs:
- Dashboard: Live RPM, error rate, P50/P95 latency, per-backend cards, and a filterable live request feed.
- Request Log: Historical request log with filters (backend, status, key, date range), paginated, with per-request cost and token detail.
- Settings: Mutable config (log level, log_bodies, per-backend model mappings), read-only env vars (secrets masked), and Export .env to generate a
.anyllm.envtemplate. - Backends: Configured backends and their settings.
- Access Control: Virtual key CRUD — create, edit (RPM/TPM limits, budget, expiry, model allowlist), and revoke keys without restarting.
- Models: Add/remove model routing deployments (LiteLLM config mode only).
- Audit: Log of all admin config mutations and key lifecycle events.
anyllm_proxy --webui
# Proxy API: http://localhost:3000
# Admin UI: http://127.0.0.1:3001/admin/?token=$(cat .admin_token)
The dashboard's Settings tab shows all active environment variables (API keys masked) and has an Export .env button that generates a .anyllm.env template you can edit and reuse. To use a custom port or a fixed token:
ADMIN_PORT=4000 ADMIN_TOKEN=mysecret anyllm_proxy --webui
To force-disable the admin even when the flag is present (useful in automated environments):
DISABLE_ADMIN=1 anyllm_proxy --webui # admin will NOT start
Additional admin env vars: ADMIN_DB_PATH (SQLite file, default: admin.db), ADMIN_TOKEN_PATH (where the generated token is written, default: .admin_token), ADMIN_LOG_RETENTION_DAYS (request log retention, default: 7).
Advanced Mode
Multiple backends on one proxy (recommended)
A single proxy instance can serve all your backends simultaneously. Each backend gets its own URL path. Use a config.toml (see section 2):
# config.toml
listen_port = 3000
default_backend = "local"
[backends.local]
kind = "openai"
api_key = "unused"
base_url = "http://localhost:11434/v1"
big_model = "qwen2.5-coder:32b"
small_model = "qwen2.5-coder:7b"
[backends.openai]
kind = "openai"
api_key = "sk-..."
base_url = "https://api.openai.com/v1"
big_model = "gpt-4o"
small_model = "gpt-4o-mini"
[backends.deepseek]
kind = "openai"
api_key = "sk-deepseek-..."
base_url = "https://api.deepseek.com/v1"
big_model = "deepseek-coder"
small_model = "deepseek-chat"
PROXY_CONFIG=config.toml anyllm_proxy --webui
All three backends are live at once:
| Path | Backend |
|---|---|
http://localhost:3000/v1/messages |
local (default) |
http://localhost:3000/openai/v1/messages |
OpenAI |
http://localhost:3000/deepseek/v1/messages |
DeepSeek |
Point different tools at different paths, or switch in Claude Code by changing ANTHROPIC_BASE_URL.
Coming from LiteLLM? Drop in your config.yaml
anyllm-proxy accepts LiteLLM config.yaml files directly. If you already have a LiteLLM deployment, point the proxy at your existing config:
PROXY_CONFIG=config.yaml anyllm_proxy --webui
A standard LiteLLM config works as-is:
# config.yaml (LiteLLM format)
model_list:
- model_name: gpt-4o
litellm_params:
model: azure/gpt-4o-eu
api_base: https://my-resource.openai.azure.com/
api_key: os.environ/AZURE_API_KEY
rpm: 6000
- model_name: gpt-4o
litellm_params:
model: openai/gpt-4o
api_key: os.environ/OPENAI_API_KEY
rpm: 10000
- model_name: claude-3-opus
litellm_params:
model: anthropic/claude-3-opus-20240229
api_key: os.environ/ANTHROPIC_API_KEY
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
Multiple deployments of the same model name are load-balanced with round-robin routing. Deployments at their RPM limit are automatically skipped.
Env var compatibility: LiteLLM env var names are accepted as aliases, so you do not need to rename anything:
| LiteLLM env var | anyllm-proxy equivalent | Notes |
|---|---|---|
LITELLM_MASTER_KEY |
PROXY_API_KEYS |
Admin/auth key |
LITELLM_CONFIG |
PROXY_CONFIG |
Config file path |
AZURE_API_KEY |
AZURE_OPENAI_API_KEY |
Azure auth |
AZURE_API_BASE |
AZURE_OPENAI_ENDPOINT |
Azure endpoint |
AZURE_API_VERSION |
AZURE_OPENAI_API_VERSION |
Azure API version |
AWS_REGION_NAME |
AWS_REGION |
Bedrock region |
OPENAI_API_KEY |
OPENAI_API_KEY |
Same name |
ANTHROPIC_API_KEY |
ANTHROPIC_API_KEY |
Same name |
The os.environ/VAR_NAME syntax in YAML values is supported alongside anyllm's native env:VAR_NAME. See docs/COMPARISON_LITELLM.md for a full feature comparison.
Multiple separate instances (for isolated deployments)
For cases where you want completely separate proxy processes (different ports, different machines, different Docker containers), keep one .env file per deployment:
~/proxies/
ollama.env # local Ollama
openai-prod.env # production OpenAI
deepseek.env # DeepSeek API
Run any one:
anyllm_proxy --env-file ~/proxies/deepseek.env
Docker-compatible — same file works with --env-file:
docker run --env-file ~/proxies/openai-prod.env -p 3000:3000 anyllm-proxy
The admin UI's Export .env button (Settings tab) generates a ready-to-edit template from the current configuration.
What, Why, and How?
What is it?
A lightweight, fast Rust-based proxy that accepts Anthropic Messages API requests, translates them to OpenAI Chat Completions format, forwards them to any compliant backend, and translates responses back in real-time. Supports streaming SSE, tool calling, and image/document blocks.
Why use it?
- Local AI Coding: Run Claude Code against local models (Llama 3, DeepSeek, Qwen) without API credits.
- Broad Compatibility: Works with open-weights and alternative models including Qwen and DeepSeek.
- Multi-Backend Routing: Route
haikurequests to a fast local model andopusrequests to external providers, transparently. - Observability: Built-in admin dashboard for request logs, latency, and live config changes.
How to Build from Source
cargo build
# Proxy only (default)
cargo run -p anyllm_proxy
# Proxy + admin web UI
cargo run -p anyllm_proxy -- --webui
The proxy listens on 0.0.0.0:3000. The admin dashboard (opt-in via --webui) binds to 127.0.0.1:3001.
1. Primary Use Case: Claude Code + Local LLMs
Claude Code works with any OpenAI-compatible backend (GPT-4o, DeepSeek, Azure, LM Studio, vLLM) by pointing it at anyllm-proxy. Use these environment variables when launching Claude Code:
export ANTHROPIC_BASE_URL="http://localhost:3000"
export ANTHROPIC_AUTH_TOKEN="proxy-user"
export ANTHROPIC_API_KEY=""
# Prevents Claude Code from hitting Anthropic's cloud telemetry endpoints
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC="1"
claude
ANTHROPIC_AUTH_TOKEN is sent as the Authorization: Bearer header to the proxy. Set it to any non-empty string (or a real virtual key if you have auth enabled). ANTHROPIC_API_KEY must be explicitly cleared so Claude Code does not attempt direct Anthropic API calls.
Example: Running with Ollama (DeepSeek / Qwen)
# 1. Start your local LLM
ollama run qwen2.5-coder:32b &
# 2. Start the translation proxy
OPENAI_API_KEY=unused \
OPENAI_BASE_URL=http://localhost:11434/v1 \
BIG_MODEL=qwen2.5-coder:32b \
SMALL_MODEL=qwen2.5-coder:32b \
cargo run -p anyllm_proxy &
# 3. Use Claude Code targeting the local proxy
export ANTHROPIC_BASE_URL="http://localhost:3000"
export ANTHROPIC_AUTH_TOKEN="proxy-user"
export ANTHROPIC_API_KEY=""
export CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC="1"
claude
Use the same pattern for LM Studio (default port 1234) or vLLM (default port 8000) by substituting OPENAI_BASE_URL.
2. Multi-Routing and the Web Interface
Create a config.toml to map different routes to different backends:
listen_port = 3000
default_backend = "local_qwen"
[backends.local_qwen]
kind = "openai"
api_key = "unused"
base_url = "http://localhost:11434/v1"
big_model = "qwen2.5-coder:32b"
small_model = "qwen2.5-coder:7b"
[backends.deepseek_api]
kind = "openai"
api_key = "sk-deepseek-..."
base_url = "https://api.deepseek.com/v1"
big_model = "deepseek-coder"
small_model = "deepseek-chat"
[backends.openrouter]
kind = "openai"
api_key = "sk-or-..."
base_url = "https://openrouter.ai/api/v1"
big_model = "anthropic/claude-3.5-sonnet"
small_model = "google/gemini-2.5-flash"
PROXY_CONFIG=config.toml anyllm_proxy --webui
Additional per-backend fields: api_format = "chat" (OpenAI only; chat or responses), omit_stream_options = true (strip stream_options for backends that reject it). Top-level log_bodies = true enables request/response body logging. Any config value can use env:VAR_NAME to read from the environment at startup (e.g., api_key = "env:OPENAI_API_KEY").
Key environment variables for advanced mode:
LOG_BODIES: Enable request/response body logging at debug level (trueor1, default: disabled).ANYLLM_DEGRADATION_WARNINGS: Set totrueor1to exposex-anyllm-degradationresponse header when translation silently drops features (default: disabled; auto-enabled whenPROXY_CONFIGis set).
All backends are live at once on a single port. The path prefix matches the backend name in the config:
| Path | Backend | Notes |
|---|---|---|
/v1/messages |
local_qwen |
default |
/deepseek_api/v1/messages |
deepseek_api |
|
/openrouter/v1/messages |
openrouter |
Point Claude Code at a specific backend:
ANTHROPIC_BASE_URL=http://localhost:3000/deepseek_api claude
The Admin Dashboard
Start the proxy with --webui, then open:
open http://127.0.0.1:3001/admin/?token=$(cat .admin_token)
The dashboard tabs are described under Admin Web Interface above. When using a LiteLLM config, the Models tab lets you add/remove deployments without editing the config file. All config mutations (model changes, key creation/revocation) are recorded in the Audit tab.
3. Commercial APIs (OpenAI, Gemini, OpenRouter)
OpenRouter:
OPENAI_API_KEY=sk-or-... \
OPENAI_BASE_URL=https://openrouter.ai/api/v1 \
BIG_MODEL=anthropic/claude-3.5-sonnet \
SMALL_MODEL=anthropic/claude-3-haiku \
cargo run -p anyllm_proxy
OpenAI:
OPENAI_API_KEY=sk-... \
BIG_MODEL=gpt-4o \
SMALL_MODEL=gpt-4o-mini \
cargo run -p anyllm_proxy
Google Gemini:
BACKEND=gemini \
GEMINI_API_KEY=AIza... \
BIG_MODEL=gemini-2.5-pro \
SMALL_MODEL=gemini-2.5-flash \
cargo run -p anyllm_proxy
4. Additional Backends and Features
Azure OpenAI
BACKEND=azure \
AZURE_OPENAI_ENDPOINT=https://myresource.openai.azure.com \
AZURE_OPENAI_DEPLOYMENT=my-gpt4o \
AZURE_OPENAI_API_KEY=... \
cargo run -p anyllm_proxy
AWS Bedrock
BACKEND=bedrock \
AWS_REGION=us-east-1 \
AWS_ACCESS_KEY_ID=AKIA... \
AWS_SECRET_ACCESS_KEY=... \
BIG_MODEL=anthropic.claude-3-5-sonnet-20241022-v2:0 \
SMALL_MODEL=anthropic.claude-3-5-haiku-20241022-v1:0 \
cargo run -p anyllm_proxy
Anthropic Passthrough
Forwards requests to the Anthropic API with no format translation. Use this when the upstream is already Anthropic and you only need auth, routing, or rate limiting from the proxy.
BACKEND=anthropic \
ANTHROPIC_API_KEY=sk-ant-... \
cargo run -p anyllm_proxy
ANTHROPIC_BASE_URL overrides the upstream URL (default: https://api.anthropic.com). Note: POST /v1/embeddings is not available on this backend.
OpenAI Chat Completions Input
The proxy accepts POST /v1/chat/completions in OpenAI format and returns OpenAI format. This means any OpenAI-native client (LiteLLM, LangChain, etc.) can route through the proxy unchanged:
curl http://localhost:3000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "x-api-key: your-key" \
-d '{
"model": "claude-sonnet-4-20250514",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 100
}'
Virtual Key Management
Create short-lived, rate-limited, or budget-capped API keys without restarting the proxy. Start with --webui to enable the admin server, then:
# Create a key with RPM/TPM limits, a monthly budget, and a model allowlist
curl -X POST http://localhost:3001/admin/api/keys \
-H "Authorization: Bearer $(cat .admin_token)" \
-H "Content-Type: application/json" \
-d '{
"description": "dev key",
"rpm_limit": 60,
"tpm_limit": 100000,
"max_budget_usd": 10.00,
"budget_duration": "monthly",
"expires_at": "2026-12-31T00:00:00Z",
"allowed_models": ["claude-*", "gpt-4o"]
}'
# Response: {"id": 1, "key": "sk-vk...", ...}
# Use the key like any other proxy key
curl http://localhost:3000/v1/messages \
-H "x-api-key: sk-vk..." \
-d '{"model": "claude-sonnet-4-20250514", "max_tokens": 100, "messages": [...]}'
# Update limits on an existing key (no restart needed)
curl -X PUT http://localhost:3001/admin/api/keys/1 \
-H "Authorization: Bearer $(cat .admin_token)" \
-H "Content-Type: application/json" \
-d '{"rpm_limit": 120, "max_budget_usd": 20.00}'
# Check spend for a key
curl http://localhost:3001/admin/api/keys/1/spend \
-H "Authorization: Bearer $(cat .admin_token)"
# Revoke immediately (no restart needed)
curl -X DELETE http://localhost:3001/admin/api/keys/1 \
-H "Authorization: Bearer $(cat .admin_token)"
budget_duration accepts daily, monthly, or lifetime. allowed_models supports exact names and prefix/* wildcards. A key at 100% of its budget returns 429 with period reset information. Webhook notifications fire at 80%, 95%, and 100% of the budget via WEBHOOK_URLS.
Requests from unauthenticated clients are rejected by default. For local development, set PROXY_OPEN_RELAY=true to accept any non-empty key (insecure, never use in production).
Distributed rate limiting (optional): Build with --features redis and set REDIS_URL=redis://localhost:6379 to use Redis-backed rate limiting across multiple proxy instances. In-process rate limits are per-instance only. RATE_LIMIT_FAIL_POLICY=open (default) allows requests when Redis is unavailable; closed rejects them with 503.
OpenTelemetry Export
cargo build -p anyllm_proxy --features otel
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 \
OTEL_SERVICE_NAME=anyllm-proxy \
OPENAI_API_KEY=sk-... \
./target/debug/anyllm_proxy
Spans are exported via OTLP HTTP (protobuf). Standard OpenTelemetry SDK environment variables control endpoint, service name, and sampling. The feature adds zero runtime overhead when not compiled in.
Using as a Library
The translation engine is available as standalone Rust crates.
crates/translator (lib, IO-free pure translation)
|
crates/client (lib, async HTTP client wrapping translator)
|
crates/proxy (bin, full proxy server)
| Level | Crate | Use Case |
|---|---|---|
| Pure translation | anyllm_translate |
Stateless type conversion between Anthropic and OpenAI formats. No IO, no HTTP. Bring your own transport. |
| HTTP client | anyllm_client |
client.messages(req).await -- send Anthropic requests, get Anthropic responses. Handles translation, HTTP, retry, and streaming internally. |
| Embedded middleware | anyllm_translate with middleware feature |
Drop-in axum Router that adds /v1/messages to an existing server. |
| Full proxy | anyllm_proxy |
Multi-backend routing, admin UI, metrics, auth. Everything in this README. |
Adding as a dependency
[dependencies]
# HTTP client (includes translation)
anyllm_client = { git = "https://github.com/whit3rabbit/anyllm-proxy" }
# Translation only (no HTTP, no async)
anyllm_translate = { git = "https://github.com/whit3rabbit/anyllm-proxy" }
# With axum middleware support
anyllm_translate = { git = "https://github.com/whit3rabbit/anyllm-proxy", features = ["middleware"] }
HTTP Client (translation + transport)
The simplest path. Send Anthropic requests, get Anthropic responses. Translation, retry, and SSE streaming are handled internally.
use anyllm_client::{Client, ClientError};
use anyllm_translate::anthropic::MessageCreateRequest;
let client = Client::builder()
.base_url("https://api.openai.com/v1/chat/completions")
.api_key("sk-...")
.build()?;
let req: MessageCreateRequest = serde_json::from_str(r#"{
"model": "claude-sonnet-4-6",
"max_tokens": 256,
"messages": [{"role": "user", "content": "Hello"}]
}"#)?;
let response = client.messages(&req).await?;
For custom TLS, SSRF protection, or per-model mapping, use ClientConfig::builder():
use anyllm_client::{Client, ClientConfig, Auth};
use anyllm_translate::TranslationConfig;
let client = Client::new(
ClientConfig::builder()
.backend_url("https://api.openai.com/v1/chat/completions")
.auth(Auth::Bearer("sk-...".into()))
.translation(
TranslationConfig::builder()
.model_map("claude-sonnet-4-6", "gpt-4o")
.model_map("claude-haiku-4-5", "gpt-4o-mini")
.build()
)
.build()
);
Error handling:
match client.messages(&req).await {
Ok(resp) => { /* ... */ }
Err(ClientError::ApiError { status, body, .. }) => eprintln!("HTTP {status}: {body}"),
Err(ClientError::Transport(e)) => eprintln!("network: {e}"),
Err(ClientError::Translation(e)) => eprintln!("translation: {e}"),
Err(e) => eprintln!("{e}"),
}
Streaming:
use anyllm_translate::anthropic::{Delta, StreamEvent};
use futures::StreamExt;
let (mut stream, _rate_limits) = client.messages_stream(&req).await?;
while let Some(event) = stream.next().await {
if let StreamEvent::ContentBlockDelta { delta: Delta::TextDelta { text }, .. } = event? {
print!("{text}");
}
}
Tool calling:
use anyllm_client::{ToolBuilder, ToolChoiceBuilder};
use serde_json::json;
let tool = ToolBuilder::new("get_weather")
.description("Get the current weather for a location")
.input_schema(json!({
"type": "object",
"properties": {"location": {"type": "string"}},
"required": ["location"]
}))
.build();
// Attach tool to MessageCreateRequest via serde_json, then call client.messages().
Runnable examples: cargo run --example basic -p anyllm_client, streaming, tools.
Pure Translation (no IO)
Use when you want to bring your own HTTP client or embed translation in a non-async context.
use anyllm_translate::{TranslationConfig, translate_request, translate_response};
use anyllm_translate::anthropic::MessageCreateRequest;
let config = TranslationConfig::builder()
.model_map("claude-sonnet-4-6", "gpt-4o")
.build();
let anthropic_req: MessageCreateRequest = serde_json::from_str(&body)?;
let openai_req = translate_request(&anthropic_req, &config)?;
// ... send openai_req with your HTTP client ...
let anthropic_resp = translate_response(&openai_resp, &anthropic_req.model);
Streaming (OpenAI chunks → Anthropic SSE events):
use anyllm_translate::new_stream_translator;
let mut translator = new_stream_translator(model);
// Feed each OpenAI chunk as it arrives:
let events = translator.process_chunk(&chunk);
// After the stream ends:
let final_events = translator.finish();
Reverse direction (OpenAI ← Anthropic), for serving OpenAI-native clients:
use anyllm_translate::{
translate_openai_to_anthropic_request,
translate_anthropic_to_openai_response,
new_reverse_stream_translator,
TranslationWarnings,
};
let mut warnings = TranslationWarnings::default();
let anthropic_req = translate_openai_to_anthropic_request(&openai_req, &mut warnings)?;
// ... forward to Anthropic API ...
let openai_resp = translate_anthropic_to_openai_response(&anthropic_resp, "gpt-4o");
Runnable examples: cargo run --example translate_request -p anyllm_translate, reverse_translation.
Embedded Middleware (for existing axum apps)
use anyllm_translate::middleware::{anthropic_compat_router, AnthropicCompatConfig};
let config = AnthropicCompatConfig::builder()
.backend_url("https://api.openai.com")
.api_key("sk-...")
.build();
let app = Router::new()
.merge(anthropic_compat_router(config))
.route("/my-other-endpoint", get(handler));
For cross-language bindings (FFI, WASM, PyO3), see docs/library-integration.md.
Advanced Features
- Streaming SSE: Real-time translation of chunked responses.
- Tool Calling: Transparent tool definition and
tool_use/tool_resulttranslation. - Image & Document Blocks: Base64/URL and document block support.
- Embeddings passthrough:
POST /v1/embeddingsforwarded as-is to the backend (no translation). Works with OpenAI, Azure, Vertex, Gemini, and vLLM. Not available whenBACKEND=anthropic. - Degradation header:
x-anyllm-degradationis set on responses when features are silently dropped during translation (e.g.,top_k,cache_control,document_blocks,thinking_config). - Model allowlist: Per-virtual-key restriction by exact model name or
prefix/*wildcard, enforced pre-request. - Budget tracking and spend alerts: Per-key
max_budget_usdwith daily/monthly/lifetime periods. Webhook notifications (viaWEBHOOK_URLS) fire at 80%, 95%, and 100% of the budget. - Audit log: All admin config mutations and key lifecycle events stored in SQLite, queryable via
GET /admin/api/audit. - OIDC/JWT authentication: Set
OIDC_ISSUER_URL(and optionallyOIDC_AUDIENCE) to accept JWT bearer tokens for proxy authentication. - Observability: SQLite request logging, metrics endpoint, WebSocket live dashboard.
- Safety: SSRF protection (including IPv6 ULA/link-local), concurrency limits, exponential backoff retry, CSRF protection on admin endpoints.
License
MIT