docs: add Phase 1 implementation plan and AGENTS.md

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
whit3rabbit
2026-03-27 06:53:24 -05:00
co-authored by Claude Opus 4.6
parent c84c2c0a68
commit 727083b2ad
2 changed files with 1278 additions and 0 deletions
+123
View File
@@ -0,0 +1,123 @@
# AGENTS.md
This file provides guidance to Codex (Codex.ai/code) when working with code in this repository.
## What This Is
An Anthropic-to-OpenAI API translation proxy in Rust. Accepts Anthropic Messages API requests, translates them to OpenAI Chat Completions format, forwards to OpenAI, and translates back. Supports streaming SSE, tool calling, file/document blocks.
See PLAN.md for the full specification and TASKS.md for phased implementation status (all 11 phases complete).
## Current Status
**Working (verified):**
- Build: `cargo build` clean, `cargo clippy -- -D warnings` clean
- Tests: ~371 tests passing (273 translator, 98 proxy)
- Full Anthropic Messages API translation: non-streaming, streaming SSE, tool calling, file/document blocks
- Proxy middleware: health, auth, request ID, size limits, concurrency limits, retry with backoff
- Compatibility endpoints: /v1/models, count_tokens (approximate via tiktoken), batches (stub)
- Model mapping and lossy-translation warnings
**Not fully validated:**
- OpenAI Responses API backend: wired up via `OPENAI_API_FORMAT=responses` but not tested against live API
- Live API integration tests exist (`crates/proxy/tests/live_api.rs`) but are `#[ignore]` by default; run with `OPENAI_API_KEY=sk-... cargo test --test live_api -- --ignored --test-threads=1`
- Metrics endpoint exists (GET /metrics returns JSON counters) but streaming requests only track total count, not success/error
## Build and Test
```bash
cargo build # build everything
cargo test # run all tests (~395 tests)
cargo test -p anyllm_translate # translator crate only
cargo test -p anyllm_proxy # proxy crate only
cargo test health_endpoint # single test by name
cargo clippy -- -D warnings # lint
cargo fmt --check # format check
```
Run the proxy (requires OPENAI_API_KEY):
```bash
OPENAI_API_KEY=sk-... cargo run -p anyllm_proxy
# Listens on 0.0.0.0:3000, health at GET /health
```
## Environment Variables
- `BACKEND`: Backend provider: `openai` (default), `vertex`, or `gemini`
- `OPENAI_API_KEY`: OpenAI API key (required when BACKEND=openai, empty default)
- `OPENAI_BASE_URL`: OpenAI base URL (default: `https://api.openai.com`)
- `OPENAI_API_FORMAT`: OpenAI API format: `chat` (default, Chat Completions) or `responses` (Responses API). Only relevant when BACKEND=openai.
- `LISTEN_PORT`: Server port (default: `3000`)
- `BIG_MODEL`: Backend model for sonnet/opus requests (default: `gpt-4o` for OpenAI, `gemini-2.5-pro` for Vertex/Gemini)
- `SMALL_MODEL`: Backend model for haiku requests (default: `gpt-4o-mini` for OpenAI, `gemini-2.5-flash` for Vertex/Gemini)
- `RUST_LOG`: Tracing filter (e.g., `info`, `anyllm_proxy=debug`)
- `TLS_CLIENT_CERT_P12`: Path to PKCS#12 (.p12/.pfx) client certificate for mTLS to the backend (optional)
- `TLS_CLIENT_CERT_PASSWORD`: Password to decrypt the P12 file (required if P12 is set)
- `TLS_CA_CERT`: Path to PEM-encoded CA certificate for verifying the backend server (optional)
- `VERTEX_PROJECT`: GCP project ID (required when BACKEND=vertex)
- `VERTEX_REGION`: GCP region, e.g. `us-central1` (required when BACKEND=vertex)
- `VERTEX_API_KEY`: Google API key for Vertex AI (one of VERTEX_API_KEY or GOOGLE_ACCESS_TOKEN required when BACKEND=vertex)
- `GOOGLE_ACCESS_TOKEN`: OAuth bearer token for Vertex AI (alternative to VERTEX_API_KEY)
- `GEMINI_API_KEY`: Google API key for Gemini Developer API (required when BACKEND=gemini)
- `GEMINI_BASE_URL`: Gemini API base URL (default: `https://generativelanguage.googleapis.com/v1beta`)
- `LOG_BODIES`: Enable request/response body logging at debug level (`true` or `1`, default: disabled)
## Architecture
Cargo workspace with two crates:
### `crates/translator` (lib: `anyllm_translate`)
Pure translation logic, no IO. Four modules:
- **`anthropic/`**: Anthropic Messages API types (request, response, streaming events, errors)
- **`openai/`**: OpenAI types for both Chat Completions and Responses APIs
- **`mapping/`**: Stateless conversion functions between the two APIs
- `message_map`: Message/content block translation (system prompt -> developer role)
- `tools_map`: Tool definitions and tool_use/tool_call translation
- `usage_map`: Token usage field mapping
- `errors_map`: HTTP status and error shape translation
- `streaming_map`: SSE event stream translation state machine
- `responses_message_map`: Anthropic to/from OpenAI Responses API mapping
- `responses_streaming_map`: Responses API SSE event stream translation state machine
- **`util/`**: JSON helpers, ID generation (uuid v4), secret redaction
### `crates/proxy` (bin: `anyllm_proxy`)
HTTP proxy built on axum + reqwest:
- **`config.rs`**: Env-based configuration
- **`server/routes.rs`**: Axum router (POST /v1/messages, GET /health, GET /metrics, GET /v1/models, stubs for count_tokens and batches)
- **`server/middleware.rs`**: Auth validation (x-api-key), request ID injection, 32MB size limit, concurrency limit, logging
- **`server/sse.rs`**: SSE response helpers for Anthropic-format streaming
- **`backend/mod.rs`**: `BackendClient` enum (OpenAI/OpenAIResponses/Vertex/Gemini), `BackendError`, shared retry helpers
- **`backend/openai_client.rs`**: reqwest client calling OpenAI Chat Completions with retry/backoff on 429/5xx
- **`backend/gemini_client.rs`**: reqwest client calling Gemini native `generateContent`/`streamGenerateContent` with retry/backoff
- **`metrics/`**: Request count, success/error tracking, exposed via GET /metrics
### Data Flow
```
Client (Anthropic format) -> proxy (axum)
-> translator: anthropic types -> mapping -> openai types
-> backend: reqwest -> OpenAI Chat Completions
-> translator: openai types -> mapping -> anthropic types
-> proxy (axum) -> Client (Anthropic format)
```
## Key Design Decisions
- The translator crate is deliberately IO-free: all mapping is pure `fn(A) -> B`. This makes it testable without mocks.
- Tool call IDs pass through directly (Anthropic tool_use.id = OpenAI tool_call.id).
- OpenAI `arguments` is a JSON string; Anthropic `input` is a JSON object. The mapping layer handles serialization.
- Streaming uses a state machine in `streaming_map.rs` that transforms OpenAI chunk events into Anthropic SSE events, with bounded channel (32) for backpressure.
- JSON fixtures in `fixtures/anthropic/` and `fixtures/openai/` are used for golden-file testing (4 fixture files).
- Retry logic: 3 retries with exponential backoff + 25% jitter, respects retry-after header.
- Backoff jitter is deterministic (upper bound, not random) to keep tests predictable.
- `ChatCompletionRequest` uses `#[serde(flatten)] pub extra: serde_json::Map` to capture unknown OpenAI fields (e.g., `seed`, `logprobs`, `logit_bias`, `n`, `reasoning_effort`). These pass through to OpenAI without typed handling. Only fields that require translation logic (not just forwarding) need explicit struct fields.
## Conventions
- Most source files reference their PLAN.md line ranges in a comment at the top.
- Test files live alongside source (`#[cfg(test)]` modules) and in `crates/proxy/tests/` for integration tests.
- Error types use `thiserror` derive macros.
- Test distribution: translator (~224 tests), proxy (~79 tests including integration/compatibility).
## References
- OpenAI API spec: https://github.com/openai/openai-openapi/blob/manual_spec/openapi.yaml (very large, ~70k+ lines). See https://simonwillison.net/2024/Dec/22/openai-openapi/ for context on the spec's size and structure. Do not attempt to load the full spec into context; reference specific sections as needed.
File diff suppressed because it is too large Load Diff