Files
anyllm-proxy/docs/gemini-api-diffs.md
T
whit3rabbitandClaude Opus 4.6 d37d1e25f1 Add phases 12-20: release infra, transparent proxy, model mapping, mTLS, extended thinking, Gemini research
Phases 12-18: release infrastructure (LICENSE, README, Dockerfile, CI,
CHANGELOG), transparent proxy with anthropic-version/anthropic-beta header
passthrough and lossy translation warnings, BIG_MODEL/SMALL_MODEL env-based
model mapping, mTLS client cert support (P12/PEM), max_completion_tokens
and reasoning_effort passthrough via serde flatten, extended thinking type
support (thinking blocks stripped in translation), top_k typed field.

Phase 20: Gemini backend research with docs/gemini-api-diffs.md covering
native API format, tool calling, streaming, auth, schema restrictions,
and Vertex AI OpenAI-compatible endpoint. Task roadmap through Phase 22.

Test count: 169 -> expanded with new fixture and unit tests for all phases.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-21 20:06:50 -05:00

8.3 KiB

Gemini API Differences from OpenAI

Research deliverable for Phase 20. Documents how Google's Gemini APIs differ from OpenAI's Chat Completions API and what translation work is required.

Recommendation

Implement Vertex OpenAI-compatible mode first (Phase 20a). Vertex AI exposes projects.locations.endpoints.openapi with Chat Completions-compatible endpoints. The existing Anthropic-to-OpenAI translation layer works as-is with different auth, base URL, and model names.

Build native Gemini translation (Phases 20c-20g) only if Vertex OpenAI-compatible mode has gaps in tool calling, streaming, or schema handling.

Two API Surfaces

Google exposes Gemini through two surfaces:

Surface Base URL Auth Model format
Gemini Developer API generativelanguage.googleapis.com/v1beta API key (x-goog-api-key header or ?key= param) models/{model}
Vertex AI {region}-aiplatform.googleapis.com OAuth bearer token (ADC / service account) projects/{p}/locations/{l}/publishers/google/models/{m}

Vertex also has an OpenAI-compatible endpoint at projects.locations.endpoints.openapi that accepts Chat Completions format directly.

Endpoint Mapping

Capability OpenAI Gemini Dev API Vertex AI
Non-streaming POST /chat/completions POST /{model}:generateContent POST .../models/{model}:generateContent
Streaming SSE from /chat/completions POST /{model}:streamGenerateContent?alt=sse (SSE) Stream of GenerateContentResponse
Embeddings POST /embeddings POST /{model}:embedContent POST .../models/*:embedContent
Token counting N/A POST /{model}:countTokens POST .../models/*:countTokens
Batch JSONL file + batch job :batchGenerateContent Vertex batch prediction

Content Model

OpenAI uses messages[] with role + content. Gemini uses contents[] with parts[].

OpenAI message:

{"role": "user", "content": "Hello"}

Gemini content:

{"role": "user", "parts": [{"text": "Hello"}]}

Parts Union

A Gemini Part can carry:

  • text: plain text
  • inlineData: {mimeType, data} (base64 bytes)
  • fileData: {mimeType, fileUri} (GCS URI for Vertex, Files API URI for Dev)
  • functionCall: {name, args} (tool invocation)
  • functionResponse: {name, response} (tool result)

OpenAI uses separate message types or structured content arrays for equivalent functionality.

Roles and System Instructions

Concept OpenAI Gemini
System prompt system role in messages systemInstruction field (separate from contents[])
Developer instructions developer role No equivalent (fold into systemInstruction)
User user user
Assistant assistant model

Gemini Content.role is restricted to user or model. The systemInstruction field only allows text parts.

Turn alternation: Gemini may enforce strict user/model alternation. Consecutive same-role messages must be merged.

Translation: Anthropic system -> Gemini systemInstruction.parts[].text. Anthropic assistant -> model. If both system and developer-style instructions exist, concatenate into systemInstruction.

Tool Calling

Aspect OpenAI Gemini
Declaration tools[].function.parameters (JSON Schema) tools[].function_declarations[].parameters (OpenAPI 3.0 subset)
Max declarations Not documented hard limit 128 per request
Call format tool_calls[].function.{name, arguments} (arguments is JSON string) Part.functionCall.{name, args} (args is JSON object)
Result format Tool role message with tool_call_id Part.functionResponse.{name, response}
Streaming calls Delta-based assembly of arguments string partialArgs[] with willContinue flag
Choice tool_choice: auto/none/required/{function} toolConfig.functionCallingConfig.mode: AUTO/NONE/ANY

Schema Restrictions

Gemini rejects certain JSON Schema constructs. Tool parameter schemas must be sanitized:

Must strip: $schema, $ref, $defs, definitions, default, pattern, examples

Must rewrite: anyOf/oneOf -> flatten to first variant or string fallback

Must validate: Only Gemini-supported format values (e.g., int32, float, date-time)

This is critical: the claude-code-proxy project implements clean_gemini_schema() for exactly this reason.

Generation Parameters

Parameter OpenAI Gemini
Max output max_tokens / max_completion_tokens maxOutputTokens
Temperature temperature temperature (range (0.0, 2.0])
Top-p top_p topP
Top-k Not always present topK
Candidates n candidateCount
Stop sequences stop (string or array) stopSequences (array)
Seed seed seed ("mostly deterministic")
Presence penalty presence_penalty presencePenalty
Frequency penalty frequency_penalty frequencyPenalty
JSON output response_format responseMimeType + responseSchema

Streaming

Gemini Developer API

Uses SSE (streamGenerateContent?alt=sse). Each SSE event contains a full GenerateContentResponse with incremental content. Parse with standard SSE framing (data: lines separated by \n\n).

Vertex AI

Documented as a "stream of GenerateContentResponse instances." Not explicitly described as SSE in REST reference. May use different HTTP framing.

OpenAI

SSE with typed chunk objects. Each chunk has choices[].delta with incremental content.

Key difference: OpenAI streams deltas (partial content). Gemini streams full response objects with accumulated content. The streaming state machine must account for this (Gemini: take last part, OpenAI: accumulate deltas).

Authentication

Surface Mechanism Header/Param
OpenAI API key Authorization: Bearer sk-...
Gemini Dev API API key x-goog-api-key: ... or ?key=...
Vertex AI OAuth token Authorization: Bearer $(gcloud auth print-access-token)
Gemini Live WS API key ?key=... in WebSocket URL

Model Naming

Claude model names do not exist in Google's ecosystem. Mapping is required:

Pattern Target
Contains "haiku" gemini-2.5-flash (small/fast)
Contains "sonnet" gemini-2.5-pro (balanced)
Contains "opus" gemini-2.5-pro (best available)
Unrecognized Passthrough with warning

Phase 15 (configurable model mapping) builds the infrastructure for this.

Safety Settings

Gemini/Vertex supports per-request safetySettings[] with {category, threshold}. Categories: hate speech, dangerous content, harassment, sexually explicit. Thresholds: BLOCK_LOW_AND_ABOVE, BLOCK_MEDIUM_AND_ABOVE, BLOCK_ONLY_HIGH, BLOCK_NONE, OFF.

No direct OpenAI equivalent. Safety rejections should map to a normalized error category, not a generic 400.

File Uploads

Feature OpenAI Gemini Dev API Vertex AI
Upload Uploads API (64MB parts, 8GB total) Files API (Google resumable upload) GCS URIs
Reference File ID in messages file.uri in fileData part GCS URI in fileData.fileUri
Headers Standard multipart X-Goog-Upload-Protocol: resumable, X-Goog-Upload-Command N/A (use gsutil/GCS client)

Error Handling

Gemini errors are not fully documented in the same structured way as OpenAI (error.message/type/param/code). Rate limit errors trigger on RPM, TPM (input), and RPD dimensions (OpenAI has 5 dimensions: RPM, RPD, TPM, TPD, IPM).

Rate Limits

Dimension OpenAI Gemini Dev API
RPM Yes Yes
RPD Yes Yes
TPM Yes Yes (input only)
TPD Yes Not documented
IPM Yes (images) Not documented

Batch Processing

OpenAI: JSONL input file (up to 50,000 requests, 200MB), uploaded with purpose: batch.

Gemini: batchGenerateContent with inlined requests or file-backed. Outputs JSONL responses preserving order.

References

  • Gemini Developer API: generativelanguage.googleapis.com/v1beta
  • Vertex AI Gemini: {region}-aiplatform.googleapis.com
  • claude-code-proxy schema sanitizer: clean_gemini_schema() pattern for JSON Schema subset restrictions
  • claude-code-via-antigravity: Gemini backend with role mapping and tool schema stripping