BerriAI/litellm
Python SDK and FastAPI proxy that put 100+ LLM providers behind OpenAI-style APIs, with routing, virtual keys, budgets and guardrails.
Overview
LiteLLM is two products in one repository. The SDK (litellm/) is a Python library: litellm.completion(model="anthropic/claude-...", messages=[...]) calls any of 100+ providers with OpenAI-style arguments and returns an OpenAI-shaped ModelResponse. The AI Gateway (litellm/proxy/) is a FastAPI server built on that SDK. It adds virtual keys, a user/team/organisation hierarchy, budgets and rate limits, spend logs in PostgreSQL, caching, guardrails, many logging integrations and a Next.js admin UI. Clients talk to it with the OpenAI, Anthropic or Azure SDKs.
The GitHub description now says “Rust core with Python SDK”. At this SHA that is a direction more than a fact. A Cargo workspace (litellm-rust/, 67 crates) is compiled into the Python wheel as litellm.rust_bridge._native, and a rollout table decides route by route whether Rust or Python runs. Chat completions, embeddings, Responses and the token counter are still Python-only. Anthropic /v1/messages can opt in to Rust. Only OCR and Bedrock transcription require it. A standalone Rust gateway binary exists, but its own README lists guardrails and spend accounting as not implemented. In practice, the production gateway is the Python proxy.
LiteLLM suits platform teams who want one OpenAI-compatible front door with cost attribution and policy, and Python developers who want a provider-neutral client. It is by far the largest of the gateways in this category, and that size cuts both ways.
Architecture
flowchart LR
C["Client: OpenAI / Anthropic / Azure SDK"] --> P["FastAPI proxy_server.py"]
P --> AU["user_api_key_auth"]
AU --> DC["DualCache: memory + Redis"]
DC -.-> PG["PostgreSQL via Prisma"]
P --> PRE["pre_call_hook: limits, guardrails"]
PRE --> RR["route_request"]
RR --> RT["Router: strategy, fallbacks, cooldowns"]
RT --> SDK["litellm.acompletion (main.py)"]
SDK --> RB["rust_bridge rollout check"]
SDK --> HH["BaseLLMHTTPHandler"]
HH --> CFG["Provider Config: transform_request / transform_response"]
CFG --> UP["Provider API"]
SDK --> LOG["Logging: cost, callbacks, spend writer"]
LOG --> PG
| Component | Path | Role |
|---|---|---|
| Proxy app | litellm/proxy/proxy_server.py |
FastAPI routes for chat, embeddings, images, Responses, files, batches and management; startup and background jobs |
| Auth | litellm/proxy/auth/ |
Virtual keys, JWT, SSO, model and route permission checks |
| Request processing | litellm/proxy/common_request_processing.py, route_llm_request.py |
Shared pre-call logic, hooks, routing to Router or SDK, response headers |
| Router | litellm/router.py, litellm/router_strategy/ |
Deployment selection, retries, fallbacks, cooldowns, routing strategies |
| SDK entry | litellm/main.py, litellm/utils.py |
completion/acompletion and provider resolution and dispatch |
| Provider layer | litellm/llms/*, llms/custom_httpx/llm_http_handler.py |
About 150 provider folders with Config transformations driven by one HTTP handler |
| Cost | litellm/cost_calculator.py, model_prices_and_context_window.json |
Per-token and per-modality prices for about 4,500 model entries |
| Caching | litellm/caching/ |
Memory, Redis, Redis Cluster, disk, S3, GCS, Azure Blob; Redis, Qdrant and Valkey semantic caches |
| Guardrails | litellm/proxy/guardrails/ |
Registry plus about 60 hook integrations (Presidio, Lakera, Bedrock, and more) |
| Rust bridge | litellm/rust_bridge/, litellm-rust/ |
Native extension, rollout rules, fallback to Python |
| Admin UI | ui/litellm-dashboard/ |
Next.js dashboard for keys, teams, models, spend and guardrails |
How a request flows
Take POST /v1/chat/completions with a virtual key:
- Endpoint.
chat_completionis registered on/v1/chat/completions,/chat/completionsand two Azure-style paths, all withDepends(user_api_key_auth). It reads the body, stamps user, team and org ids intometadata, and hands off toProxyBaseLLMRequestProcessing.base_process_llm_request(proxy_server.py). - Authenticate.
user_api_key_authaccepts the key fromAuthorization,api-key(Azure),x-api-key(Anthropic), the Google header or a custom header. It resolves the key’s identity through the cache or the database inside an OTelauthspan, then authorises the model and route (user_api_key_auth.py). - Pre-call hooks.
proxy_logging_obj.pre_call_hookruns the proxy hooks (parallel-request and TPM/RPM limiter, cache control, budgets) and pre-call guardrails, which may mask or block the input (common_request_processing.py). - Route.
during_call_hook(moderation-style guardrails) is started as a task in parallel with the LLM call.route_requestthen resolves the model (common_request_processing.py). Team model aliases, router model groups, wildcards and deployment names go tollm_router.acompletion. Withpass_through_all_models, unknown models go straight tolitellm.acompletion(route_llm_request.py). - Router.
Router.acompletiongoes throughasync_function_with_fallbacks, which wraps retries and then the configured fallbacks, context-window fallbacks and content-policy fallbacks (router.py)._acompletionpicks a deployment withasync_get_available_deployment, may mirror the call to a “silent” experiment model, takes a per-deploymentmax_parallel_requestsslot, compacts the context if needed, and callslitellm.acompletion(router.py). - SDK dispatch.
completion()resolves the provider and walks a longif custom_llm_provider == ...chain. Most providers end inbase_llm_http_handler.completion(main.py, main.py). The handler calls the providerConfig’stransform_request, sends the HTTP request and callstransform_responseor the streaming iterator. - Cost and logging. The logging object computes
response_costwithcompletion_cost, puts it in_hidden_params, and the proxy returns it asx-litellm-response-costheaders. Success callbacks queue spend increments, which a background job flushes to PostgreSQL about every 60 seconds (ARCHITECTURE.md).
Key components
Router and strategies
routing_strategy_init builds one selector per strategy and supports per-model-group routing_groups with their own strategy (router.py). The default simple-shuffle draws by weight, then rpm, then tpm, from healthy deployments (simple_shuffle.py). Other strategies in router_strategy/ are least-busy, lowest-latency, lowest-cost, usage-based (TPM/RPM), tag-based, budget-limited, and newer complexity, quality, adaptive and auto routers. Failed deployments enter a cooldown controlled by allowed_fails and cooldown_time. Streams can fall back mid-response.
Provider translation
Each provider has a Config class that extends BaseConfig and implements the abstract transform_request and transform_response (transformation.py). BaseLLMHTTPHandler owns the HTTP, retry and streaming mechanics, so adding a provider rarely touches the handler. The older BaseLLM class in llms/base.py is a legacy template. Native-format endpoints exist too: /v1/messages translates Anthropic Messages to any backend, and pass-through routes forward raw provider APIs with LiteLLM auth and spend tracking.
Keys, budgets and spend
The Prisma schema has 94 models, including LiteLLM_VerificationToken, LiteLLM_TeamTable, LiteLLM_UserTable, organisations, budgets and spend logs. Keys are hashed. Budgets, TPM, RPM and parallel-request limits can be set on keys, users, teams, organisations, end users and tags. Spend is accumulated in Redis or in memory and written in batches, so a burst of requests does not become a burst of database writes.
Caching
Cache takes a backend type, a mode (default_on or default_off), TTLs and a similarity_threshold for semantic backends (caching.py). DualCache (memory plus Redis) is also the internal store for key lookups, rate-limit counters and cooldowns.
Guardrails and callbacks
Guardrails subclass CustomGuardrail, itself a CustomLogger. They can run pre-call, during the call (in parallel with the LLM) or post-call, including on stream chunks. The registry auto-discovers about 60 integrations in guardrail_hooks/: Presidio PII masking, Lakera, Bedrock Guardrails, Model Armor, Pangea, a built-in content filter, tool permission policies and more. The same callback system feeds Langfuse, Datadog, OpenTelemetry, Prometheus and many other loggers.
Rust bridge
catalog.RULES is the rollout table (catalog.py). RUST_OPT_IN routes switch on with LITELLM_RUST=true or litellm.rust(True). They fall back to Python only if the native side declines before execution starts (configuration.py, dispatch.py). The separate Axum gateway in litellm-rust/crates/gateway serves inference and MCP routes from a YAML config. It does not yet implement guardrails, spend accounting or database-backed management (README).
Extending it
- Provider. Add
llms/<provider>/chat/transformation.pywith aBaseConfigsubclass. Add pricing entries tomodel_prices_and_context_window.json. - Hooks. Implement
CustomLogger(logging, pre/post-call hooks) orCustomGuardrailand reference it from the proxy config. Proxy-wide hooks are registered inPROXY_HOOKS. - Routing. Choose a
routing_strategyper router or per routing group, or plug in a custom strategy. - Pass-through. Expose a provider’s native API through
pass_through_endpointswhile still enforcing keys and budgets. - Split deployment.
gateway/main.pyandbackend/main.pyreuse the same FastAPI app but trim it to the data-plane or the admin routes, so inference and management can scale separately (gateway/main.py).
Running it
- SDK.
pip install litellmand callcompletion(model="provider/model", ...)with provider keys in the environment. - Proxy.
pip install 'litellm[proxy]', thenlitellm --config config.yaml. The CLI launches uvicorn onlitellm.proxy.proxy_server:appwith options for workers, keep-alive, JSON logs and worker health checks (proxy_cli.py). - Containers. The Docker image starts on port 4000.
docker-compose.ymladds PostgreSQL and Prometheus. Helm charts and Terraform modules live inhelm/andterraform/. Keys, teams, the UI and spend tracking needDATABASE_URLand aLITELLM_MASTER_KEY. Redis is needed for shared limits and cooldowns across instances. - Building from source compiles the Rust extension through maturin, so a Rust toolchain is required.
Strengths and caveats
- Strength: breadth. About 150 provider folders, OpenAI, Anthropic and Responses-style endpoints, batches, files, vector stores, MCP and pass-through make it the most complete gateway here.
- Strength: cost attribution. A maintained price table, per-request cost headers and hierarchical budgets are hard to match.
- Strength: policy plumbing. One hook system covers rate limits, guardrails, logging and alerts, and almost everything is configurable without a fork.
- Caveat: size and complexity.
router.pyis over 15,000 lines andproxy_server.pyover 20,000. Behaviour depends on many interacting settings, and debugging a request means following a long path. - Caveat: Python hot path. Despite the “Rust core” label, the main chat path at this SHA is Python and asyncio. Throughput per instance comes from uvicorn workers, not native code.
- Caveat: open-core. Code under
enterprise/has a commercial licence, and a few dozen proxy features check for a premium licence. Check that the feature you need is in the MIT part. - Caveat: infrastructure. Useful multi-tenant deployments need PostgreSQL and Redis, plus a migration story for Prisma schema changes.
Sources: code at 62dee3d, verified Q&A.
How it answers the LLM gateways questions
Each answer was drafted by a code-reading agent at commit 62dee3d. Its citations were checked mechanically. Compare with the other llm gateways →
How are requests routed across providers and models?
answeredLoad balancing. The Router class (litellm/router.py) supports multiple strategies: simple-shuffle (weighted random: deployments with higher weight, rpm, or tpm config values get proportionally more traffic — litellm/router_strategy/simple_shuffle.py:43-67), least-busy (picks the deployment with fewest in-flight requests, tracked via a Redis-backed counter incremented pre-call and decremented on success/failure — litellm/router_strategy/least_busy.py:114-225), latency-based-routing (selects the deployment with the lowest average or percentile time-to-first-token over a configurable window — litellm/router_strategy/lowest_latency.py:53-80), cost-based-routing (tracks token usage per deployment in minute buckets via cost_map:{model_group} cache keys to pick the cheapest — litellm/router_strategy/lowest_cost.py:15-80), usage-based-routing-v2 (T/RPM-aware), and LAR-1 (semantic routing by agent confidence thresholds — litellm/router_strategy/lar1_routing.py:1-60). Additional strategies include complexity-based (complexity_router/), quality-based (quality_router/), and adaptive (adaptive_router/) routers.
Fallbacks and retries. The key entry point Router.async_function_with_fallbacks() (router.py:7865) wraps every provider call. On failure it checks two levels: order-based fallback (tries higher target_order deployments within the same model group) then configured external fallbacks (explicit fallbacks=[...] or per-error-type context_window_fallbacks, content_policy_fallbacks). A RetryPolicy per deployment controls retry counts per error category. Failed deployments enter cooldown (configurable allowed_fails and cooldown_time). Mid-stream fallbacks are supported via MidStreamFallbackError + FallbackAwareStreamWrapper (router.py:828-867,3061-3167).
Model aliases. model_group_alias (router.py:1119) maps alias names to deployment groups, and Router.get_model_from_alias() resolves them at routing time (router.py:5382).
Health checks. async_get_available_deployment (router.py:13711-13831) runs pre-routing hooks and checks deployment health via configurable staleness thresholds before selecting a target. The health-check subsystem (litellm/proxy/health_check.py) pings endpoints periodically.
How are different provider APIs unified?
answeredLiteLLM unifies 150+ providers under an OpenAI-compatible schema. The core insight is in litellm/__init__.py (line 1149) and 20 more like deepseek, groq, perplexity, cerebras) that speak OpenAI's wire format directly. Each provider implements a subclass of litellm/constants.py (line ~691): LITELLM_CHAT_PROVIDERS lists ~70 native chat providers (openai, anthropic, gemini, cohere, together_ai, etc.) plus openai_compatible_providers (BaseLLM (litellm/llms/base.py:14-80) that overrides process_response() and validate_environment().
Request translation. Each provider subdirectory under litellm/llms/ (Anthropic, Bedrock, Gemini, Cohere, etc.) contains a provider module that translates the incoming OpenAI-format request (messages, tools, stream, response_format) into the provider's native format and then maps the response back to ModelResponse — LiteLLM's pydantic model that mirrors openai.types.completion.ChatCompletion. For example litellm/llms/anthropic/ handles Anthropic's different tool-call, thinking, and streaming formats.
Streaming. Each provider normalizes its stream chunks into OpenAI-format ChatCompletionChunk objects via a CustomStreamWrapper, so downstream code sees the same SSE event structure. Tool calls and multimodal (image/audio/video inputs mapped to OpenAI content-block arrays) are translated per-provider.
Cost table. The 80,644-line model_prices_and_context_window.json at the repo root is the authoritative price list with input_cost_per_token, output_cost_per_token, max_input_tokens, mode, supports_* flags for every known model — parsed at runtime by cost_calculator.py:1-80.
Pass-through routes. The proxy also supports raw provider-API pass-through for endpoints it hasn't modeled (litellm/proxy/pass_through_endpoints/), allowing the proxy to act as a plain credential/rate-limit gate for any provider API.
Config classes that extend BaseConfig (llms/base_llm/chat/transformation.py) with transform_request/transform_response, called by BaseLLMHTTPHandler; BaseLLM.process_response in llms/base.py is a legacy template.How are API keys, users and tenants managed?
answeredVirtual keys (Verification Tokens). The core auth model is LiteLLM_VerificationToken in the Prisma schema (schema.prisma), which stores hashed API keys with associated permissions (models, budgets, teams, end-users, metadata). The user_api_key_auth.py:3487+ function is the FastAPI dependency called on every proxied request — it looks up the bearer token, validates it against the DB (with Redis cache backing via UserApiKeyCache), and returns a UserAPIKeyAuth object that downstream handlers use for access decisions.
Multi-tenant hierarchy. The data model supports a full RBAC hierarchy: Organizations → Teams → Users → API Keys, with Team Memberships linking Users to Teams with role/permission scoping. Budgets cascade: an Organization budget limits its Teams, a Team budget limits its Members and Keys. The auth check in _can_object_call_model() (litellm/proxy/auth/auth_checks.py) traverses this chain.
Upstream credential storage. Provider API keys are stored in LiteLLM_CredentialsTable (schema.prisma) and can be resolved from environment variables, AWS Secrets Manager, Google Cloud KMS, HashiCorp Vault, or custom secret managers (litellm/secret_managers/).
Auth methods. The proxy accepts API keys via the Authorization: Bearer header (OpenAI-compatible), x-api-key (Anthropic-compatible), or api-key (Azure-compatible) — see litellm/proxy/auth/user_api_key_auth.py:183-199. It also supports JWT auth (litellm/proxy/auth/handle_jwt.py), OAuth2 proxy hooks, and custom SSO.
Admin UI. The /ui route serves a React-based admin dashboard. Management endpoints under /key/, /user/, /team/, /organization/ (litellm/proxy/management_endpoints/) allow CRUD for all auth entities. A master key (LITELLM_MASTER_KEY) bootstraps the first admin.
How are rate limits, budgets and cost tracking implemented?
answeredRate limits. The LiteLLM_BudgetTable in schema.prisma defines max_parallel_requests, tpm_limit (tokens per minute), rpm_limit (requests per minute), and tpd_limit (tokens per day) at every level of the hierarchy (org, team, user, key, end-user, tag). These are enforced via the budget throttle middleware (litellm/proxy/auth/budget_throttle.py) and pre-call RPM/TPM checks in the routing layer (router.py:8984-9033).
Spend and budget tracking. litellm/proxy/spend_tracking/ implements a batched spend counter (spend_counter_batch.py) that increments in-memory counters and flushes to PostgreSQL asynchronously on a cadence. The LiteLLM_SpendLog* tables in the Prisma schema record every request's token usage, cost, and metadata. Budget enforcement checks (_virtual_key_max_budget_check, _virtual_key_soft_budget_check in litellm/proxy/auth/auth_checks.py) compare accumulated spend against max_budget/soft_budget thresholds.
Price table. model_prices_and_context_window.json (80,644 lines) is the bundled pricing catalog. Every model entry specifies input_cost_per_token, output_cost_per_token, output_cost_per_reasoning_token, regional uplift multipliers, prompt-caching discounts, batch rate reductions, and modality-specific costs (audio, image, video, web-search, computer-use). The cost calculator (litellm/cost_calculator.py) uses a chain-of-responsibility pattern: each provider module registers a cost_per_token() function (e.g. openai_cost_per_token, anthropic_cost_per_token, bedrock_cost_per_token in litellm/llms/{provider}/cost_calculation.py), and the base resolver selects the right one.
Usage accounting. The final response_cost, total_tokens, and breakdown are stored in the LiteLLM_SpendLog* tables. Budget carry-forward across reset periods is handled by carried_budget_state.py. Enterprise users get PTU (pay-per-use) flat-cost pricing via ptu_flat_cost_rollup.py.
How are caching and guardrails implemented?
answeredExact response caching. The Cache class (litellm/caching/caching.py:74+) supports multiple backends: in-memory (InMemoryCache), Redis (RedisCache), Redis Cluster (RedisClusterCache), DiskCache, S3, GCS, and Azure Blob. The DualCache wrapper maintains a local in-memory LRU + a remote Redis cache for fast local reads with cluster consistency. Response caching is controlled by CacheMode (default_on vs default_off) and per-request ttl.
Semantic caching. Two vector-based semantic caches are available: QdrantSemanticCache (litellm/caching/qdrant_semantic_cache.py) embeds prompts via a configurable embedding model and queries the Qdrant vector DB for semantically similar requests, returning the cached response if similarity exceeds a threshold. RedisSemanticCache (litellm/caching/redis_semantic_cache.py) does the same using Redis Stack's vector-similarity-search (VSS) capabilities.
Guardrails. The GuardrailRegistry in litellm/proxy/guardrails/guardrail_registry.py:222-386 maintains guardrail_class_registry — a dict mapping integration names to CustomGuardrail subclasses. ~50 guardrail hooks are auto-discovered from litellm/proxy/guardrails/guardrail_hooks/ via get_guardrail_class_from_hooks() (line 310), which scans subdirectories for guardrail_class_registry dicts. Built-in integrations include:
- PII redaction: Presidio-based
_OPTIONAL_PresidioPIIMasking(guardrail_hooks/presidio.py:172+) analyzes and anonymizes PII (names, emails, SSNs, credit cards) in request/response content, including SSE stream chunks. - Moderation: Lakera AI (
guardrail_hooks/lakera_ai.py:49) for prompt injection and content moderation. Also OpenAI Moderation, Bedrock Guardrails, and Guardrails AI. - Custom hooks: Any guardrail can implement
CustomGuardrail(litellm/integrations/custom_guardrail.py) with pre-request (modify/block) and post-response callbacks, plus streaming support for SSE-based anonymization.
Plugin points. Guardrails register as standard litellm.callbacks via CustomLogger hooks, so they also participate in the general callback pipeline (pre-request, post-success, post-failure, streaming).
How is it observed, deployed and scaled?
answeredObservability (OpenTelemetry). LiteLLM has a comprehensive OTel integration. The opentelemetry_logger (litellm/integrations/otel/logger.py) implements CustomLogger and emits spans for every LLM call, routing decision, guardrail invocation, and service operation. Span attributes include request metadata, deployment info, latency, token counts, and cost. Routing decisions are captured via routing_decision_attributes() (litellm/integrations/otel/routing.py:48-57) which emits ~40 typed fields per decision. The tracing/ module (literllm/tracing/__init__.py) provides an OTLP HTTP receiver for forwarding spans to ClickHouse or any OTel-compatible backend. Prometheus metrics are exposed at /metrics (litellm/proxy/prometheus_metrics_server.py).
Architecture and concurrency. Pure Python + asyncio. The proxy is a FastAPI app deployed via uvicorn (ASGI server, configurable worker count via --num_workers, proxy_cli.py:295-329). Async at every layer: httpx.AsyncClient for outgoing provider calls, asyncio for concurrent request handling, and asyncpg via Prisma for DB access. Redis caching uses redis-py async. Deployment slots use Python anyio semaphores for max-parallel-requests enforcement (router.py:9035-9050).
Deployment modes. The CLI (litellm/proxy/proxy_cli.py) uses click to launch uvicorn with --port, --host, --num_workers, --max_requests_before_restart (for memory leak defense), --timeout_worker_healthcheck, and JSON logging support. Official Docker images (Dockerfile, docker-compose.yml) include PostgreSQL + Redis containers. Helm charts and Terraform scripts are provided for Kubernetes deployments. The --config flag points to a YAML/JSON config file defining models, routing, auth, and guardrail settings.
High availability. PostgreSQL is the durable backing store (via Prisma ORM with connection pooling via PgBouncer). Redis can be clustered for caching resilience. LiteLLM supports DB-native compaction, query-engine reaping, and graceful worker restarts. Logging is pluggable via litellm.callbacks — Langfuse, Datadog, New Relic, and custom loggers are first-class integrations.
litellm-rust/, loaded as litellm.rust_bridge._native) ships in the wheel; per-route rollout rules keep chat and embeddings on Python, let Anthropic messages opt in via LITELLM_RUST, and require Rust for OCR and Bedrock transcription.