# BerriAI/litellm

> Python SDK and FastAPI proxy that put 100+ LLM providers behind OpenAI-style APIs, with routing, virtual keys, budgets and guardrails.

- Category: [LLM gateways](https://llms-technical-reviews.com/llm-gateways/)
- Repository: https://github.com/BerriAI/litellm (reviewed at commit `62dee3d73046fb717f8693a7660d2fe457c9370f`, 2026-10-06)
- Stars: 60246 · Language: Python · License: n/a
- Canonical page: https://llms-technical-reviews.com/p/litellm/

## Overview

LiteLLM is two products in one repository. The **SDK** (`litellm/`) is a Python library: `litellm.completion(model="anthropic/claude-...", messages=[...])` calls any of 100+ providers with OpenAI-style arguments and returns an OpenAI-shaped `ModelResponse`. The **AI Gateway** (`litellm/proxy/`) is a FastAPI server built on that SDK. It adds virtual keys, a user/team/organisation hierarchy, budgets and rate limits, spend logs in PostgreSQL, caching, guardrails, many logging integrations and a Next.js admin UI. Clients talk to it with the OpenAI, Anthropic or Azure SDKs.

The GitHub description now says "Rust core with Python SDK". At this SHA that is a direction more than a fact. A Cargo workspace (`litellm-rust/`, 67 crates) is compiled into the Python wheel as `litellm.rust_bridge._native`, and a rollout table decides route by route whether Rust or Python runs. Chat completions, embeddings, Responses and the token counter are still Python-only. Anthropic `/v1/messages` can opt in to Rust. Only OCR and Bedrock transcription require it. A standalone Rust gateway binary exists, but its own README lists guardrails and spend accounting as not implemented. In practice, the production gateway is the Python proxy.

LiteLLM suits platform teams who want one OpenAI-compatible front door with cost attribution and policy, and Python developers who want a provider-neutral client. It is by far the largest of the gateways in this category, and that size cuts both ways.

## Architecture

```mermaid
flowchart LR
  C["Client: OpenAI / Anthropic / Azure SDK"] --> P["FastAPI proxy_server.py"]
  P --> AU["user_api_key_auth"]
  AU --> DC["DualCache: memory + Redis"]
  DC -.-> PG["PostgreSQL via Prisma"]
  P --> PRE["pre_call_hook: limits, guardrails"]
  PRE --> RR["route_request"]
  RR --> RT["Router: strategy, fallbacks, cooldowns"]
  RT --> SDK["litellm.acompletion (main.py)"]
  SDK --> RB["rust_bridge rollout check"]
  SDK --> HH["BaseLLMHTTPHandler"]
  HH --> CFG["Provider Config: transform_request / transform_response"]
  CFG --> UP["Provider API"]
  SDK --> LOG["Logging: cost, callbacks, spend writer"]
  LOG --> PG
```

| Component | Path | Role |
|---|---|---|
| Proxy app | `litellm/proxy/proxy_server.py` | FastAPI routes for chat, embeddings, images, Responses, files, batches and management; startup and background jobs |
| Auth | `litellm/proxy/auth/` | Virtual keys, JWT, SSO, model and route permission checks |
| Request processing | `litellm/proxy/common_request_processing.py`, `route_llm_request.py` | Shared pre-call logic, hooks, routing to Router or SDK, response headers |
| Router | `litellm/router.py`, `litellm/router_strategy/` | Deployment selection, retries, fallbacks, cooldowns, routing strategies |
| SDK entry | `litellm/main.py`, `litellm/utils.py` | `completion`/`acompletion` and provider resolution and dispatch |
| Provider layer | `litellm/llms/*`, `llms/custom_httpx/llm_http_handler.py` | About 150 provider folders with `Config` transformations driven by one HTTP handler |
| Cost | `litellm/cost_calculator.py`, `model_prices_and_context_window.json` | Per-token and per-modality prices for about 4,500 model entries |
| Caching | `litellm/caching/` | Memory, Redis, Redis Cluster, disk, S3, GCS, Azure Blob; Redis, Qdrant and Valkey semantic caches |
| Guardrails | `litellm/proxy/guardrails/` | Registry plus about 60 hook integrations (Presidio, Lakera, Bedrock, and more) |
| Rust bridge | `litellm/rust_bridge/`, `litellm-rust/` | Native extension, rollout rules, fallback to Python |
| Admin UI | `ui/litellm-dashboard/` | Next.js dashboard for keys, teams, models, spend and guardrails |

## How a request flows

Take `POST /v1/chat/completions` with a virtual key:

1. **Endpoint.** `chat_completion` is registered on `/v1/chat/completions`, `/chat/completions` and two Azure-style paths, all with `Depends(user_api_key_auth)`. It reads the body, stamps user, team and org ids into `metadata`, and hands off to `ProxyBaseLLMRequestProcessing.base_process_llm_request` ([proxy_server.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/proxy_server.py#L12045-L12135)).
2. **Authenticate.** `user_api_key_auth` accepts the key from `Authorization`, `api-key` (Azure), `x-api-key` (Anthropic), the Google header or a custom header. It resolves the key's identity through the cache or the database inside an OTel `auth` span, then authorises the model and route ([user_api_key_auth.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/auth/user_api_key_auth.py#L3487-L3545)).
3. **Pre-call hooks.** `proxy_logging_obj.pre_call_hook` runs the proxy hooks (parallel-request and TPM/RPM limiter, cache control, budgets) and pre-call guardrails, which may mask or block the input ([common_request_processing.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/common_request_processing.py#L2224-L2235)).
4. **Route.** `during_call_hook` (moderation-style guardrails) is started as a task in parallel with the LLM call. `route_request` then resolves the model ([common_request_processing.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/common_request_processing.py#L2615-L2643)). Team model aliases, router model groups, wildcards and deployment names go to `llm_router.acompletion`. With `pass_through_all_models`, unknown models go straight to `litellm.acompletion` ([route_llm_request.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/route_llm_request.py#L759-L780)).
5. **Router.** `Router.acompletion` goes through `async_function_with_fallbacks`, which wraps retries and then the configured fallbacks, context-window fallbacks and content-policy fallbacks ([router.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L7865-L7899)). `_acompletion` picks a deployment with `async_get_available_deployment`, may mirror the call to a "silent" experiment model, takes a per-deployment `max_parallel_requests` slot, compacts the context if needed, and calls `litellm.acompletion` ([router.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L3788-L3892)).
6. **SDK dispatch.** `completion()` resolves the provider and walks a long `if custom_llm_provider == ...` chain. Most providers end in `base_llm_http_handler.completion` ([main.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/main.py#L5845-L5900), [main.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/main.py#L1588-L1600)). The handler calls the provider `Config`'s `transform_request`, sends the HTTP request and calls `transform_response` or the streaming iterator.
7. **Cost and logging.** The logging object computes `response_cost` with `completion_cost`, puts it in `_hidden_params`, and the proxy returns it as `x-litellm-response-cost` headers. Success callbacks queue spend increments, which a background job flushes to PostgreSQL about every 60 seconds ([ARCHITECTURE.md](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/ARCHITECTURE.md#L217-L240)).

## Key components

### Router and strategies

`routing_strategy_init` builds one selector per strategy and supports per-model-group `routing_groups` with their own strategy ([router.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L1565-L1592)). The default `simple-shuffle` draws by `weight`, then `rpm`, then `tpm`, from healthy deployments ([simple_shuffle.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/simple_shuffle.py#L43-L60)). Other strategies in `router_strategy/` are least-busy, lowest-latency, lowest-cost, usage-based (TPM/RPM), tag-based, budget-limited, and newer complexity, quality, adaptive and auto routers. Failed deployments enter a cooldown controlled by `allowed_fails` and `cooldown_time`. Streams can fall back mid-response.

### Provider translation

Each provider has a `Config` class that extends `BaseConfig` and implements the abstract `transform_request` and `transform_response` ([transformation.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/llms/base_llm/chat/transformation.py#L313-L359)). `BaseLLMHTTPHandler` owns the HTTP, retry and streaming mechanics, so adding a provider rarely touches the handler. The older `BaseLLM` class in `llms/base.py` is a legacy template. Native-format endpoints exist too: `/v1/messages` translates Anthropic Messages to any backend, and pass-through routes forward raw provider APIs with LiteLLM auth and spend tracking.

### Keys, budgets and spend

The Prisma schema has 94 models, including `LiteLLM_VerificationToken`, `LiteLLM_TeamTable`, `LiteLLM_UserTable`, organisations, budgets and spend logs. Keys are hashed. Budgets, TPM, RPM and parallel-request limits can be set on keys, users, teams, organisations, end users and tags. Spend is accumulated in Redis or in memory and written in batches, so a burst of requests does not become a burst of database writes.

### Caching

`Cache` takes a backend type, a mode (`default_on` or `default_off`), TTLs and a `similarity_threshold` for semantic backends ([caching.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/caching/caching.py#L74-L90)). `DualCache` (memory plus Redis) is also the internal store for key lookups, rate-limit counters and cooldowns.

### Guardrails and callbacks

Guardrails subclass `CustomGuardrail`, itself a `CustomLogger`. They can run pre-call, during the call (in parallel with the LLM) or post-call, including on stream chunks. The registry auto-discovers about 60 integrations in `guardrail_hooks/`: Presidio PII masking, Lakera, Bedrock Guardrails, Model Armor, Pangea, a built-in content filter, tool permission policies and more. The same callback system feeds Langfuse, Datadog, OpenTelemetry, Prometheus and many other loggers.

### Rust bridge

`catalog.RULES` is the rollout table ([catalog.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/rust_bridge/catalog.py#L68-L79)). `RUST_OPT_IN` routes switch on with `LITELLM_RUST=true` or `litellm.rust(True)`. They fall back to Python only if the native side declines before execution starts ([configuration.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/rust_bridge/configuration.py#L47-L93), [dispatch.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/rust_bridge/dispatch.py#L80-L91)). The separate Axum gateway in `litellm-rust/crates/gateway` serves inference and MCP routes from a YAML config. It does not yet implement guardrails, spend accounting or database-backed management ([README](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm-rust/crates/gateway/README.md#L1-L43)).

## Extending it

- **Provider.** Add `llms/<provider>/chat/transformation.py` with a `BaseConfig` subclass. Add pricing entries to `model_prices_and_context_window.json`.
- **Hooks.** Implement `CustomLogger` (logging, pre/post-call hooks) or `CustomGuardrail` and reference it from the proxy config. Proxy-wide hooks are registered in `PROXY_HOOKS`.
- **Routing.** Choose a `routing_strategy` per router or per routing group, or plug in a custom strategy.
- **Pass-through.** Expose a provider's native API through `pass_through_endpoints` while still enforcing keys and budgets.
- **Split deployment.** `gateway/main.py` and `backend/main.py` reuse the same FastAPI app but trim it to the data-plane or the admin routes, so inference and management can scale separately ([gateway/main.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/gateway/main.py#L1-L60)).

## Running it

- **SDK.** `pip install litellm` and call `completion(model="provider/model", ...)` with provider keys in the environment.
- **Proxy.** `pip install 'litellm[proxy]'`, then `litellm --config config.yaml`. The CLI launches uvicorn on `litellm.proxy.proxy_server:app` with options for workers, keep-alive, JSON logs and worker health checks ([proxy_cli.py](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/proxy_cli.py#L295-L329)).
- **Containers.** The Docker image starts on port 4000. `docker-compose.yml` adds PostgreSQL and Prometheus. Helm charts and Terraform modules live in `helm/` and `terraform/`. Keys, teams, the UI and spend tracking need `DATABASE_URL` and a `LITELLM_MASTER_KEY`. Redis is needed for shared limits and cooldowns across instances.
- **Building from source** compiles the Rust extension through maturin, so a Rust toolchain is required.

## Strengths and caveats

- **Strength: breadth.** About 150 provider folders, OpenAI, Anthropic and Responses-style endpoints, batches, files, vector stores, MCP and pass-through make it the most complete gateway here.
- **Strength: cost attribution.** A maintained price table, per-request cost headers and hierarchical budgets are hard to match.
- **Strength: policy plumbing.** One hook system covers rate limits, guardrails, logging and alerts, and almost everything is configurable without a fork.
- **Caveat: size and complexity.** `router.py` is over 15,000 lines and `proxy_server.py` over 20,000. Behaviour depends on many interacting settings, and debugging a request means following a long path.
- **Caveat: Python hot path.** Despite the "Rust core" label, the main chat path at this SHA is Python and asyncio. Throughput per instance comes from uvicorn workers, not native code.
- **Caveat: open-core.** Code under `enterprise/` has a commercial licence, and a few dozen proxy features check for a premium licence. Check that the feature you need is in the MIT part.
- **Caveat: infrastructure.** Useful multi-tenant deployments need PostgreSQL and Redis, plus a migration story for Prisma schema changes.

*Sources: code at 62dee3d, verified Q&A.*

## How BerriAI/litellm answers the LLM gateways questions

### How are requests routed across providers and models? (answered)

**Load balancing.** The `Router` class (`litellm/router.py`) supports multiple strategies: **simple-shuffle** (weighted random: deployments with higher `weight`, `rpm`, or `tpm` config values get proportionally more traffic — `litellm/router_strategy/simple_shuffle.py:43-67`), **least-busy** (picks the deployment with fewest in-flight requests, tracked via a Redis-backed counter incremented pre-call and decremented on success/failure — `litellm/router_strategy/least_busy.py:114-225`), **latency-based-routing** (selects the deployment with the lowest average or percentile time-to-first-token over a configurable window — `litellm/router_strategy/lowest_latency.py:53-80`), **cost-based-routing** (tracks token usage per deployment in minute buckets via `cost_map:{model_group}` cache keys to pick the cheapest — `litellm/router_strategy/lowest_cost.py:15-80`), **usage-based-routing-v2** (T/RPM-aware), and **LAR-1** (semantic routing by agent confidence thresholds — `litellm/router_strategy/lar1_routing.py:1-60`). Additional strategies include complexity-based (`complexity_router/`), quality-based (`quality_router/`), and adaptive (`adaptive_router/`) routers.

**Fallbacks and retries.** The key entry point `Router.async_function_with_fallbacks()` (`router.py:7865`) wraps every provider call. On failure it checks two levels: order-based fallback (tries higher `target_order` deployments within the same model group) then configured external fallbacks (explicit `fallbacks=[...]` or per-error-type `context_window_fallbacks`, `content_policy_fallbacks`). A `RetryPolicy` per deployment controls retry counts per error category. Failed deployments enter cooldown (configurable `allowed_fails` and `cooldown_time`). Mid-stream fallbacks are supported via `MidStreamFallbackError` + `FallbackAwareStreamWrapper` (`router.py:828-867,3061-3167`).

**Model aliases.** `model_group_alias` (`router.py:1119`) maps alias names to deployment groups, and `Router.get_model_from_alias()` resolves them at routing time (`router.py:5382`).

**Health checks.** `async_get_available_deployment` (`router.py:13711-13831`) runs pre-routing hooks and checks deployment health via configurable staleness thresholds before selecting a target. The health-check subsystem (`litellm/proxy/health_check.py`) pings endpoints periodically.


Citations: [litellm/router_strategy/simple_shuffle.py:43-67](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/simple_shuffle.py#L43-L67) · [litellm/router_strategy/least_busy.py:114-225](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/least_busy.py#L114-L225) · [litellm/router_strategy/lowest_latency.py:53-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/lowest_latency.py#L53-L80) · [litellm/router_strategy/lowest_cost.py:15-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router_strategy/lowest_cost.py#L15-L80) · [litellm/router.py:7865-7870](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L7865-L7870) · [litellm/router.py:13711-13831](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L13711-L13831)

### How are different provider APIs unified? (answered)

LiteLLM unifies 150+ providers under an **OpenAI-compatible schema**. The core insight is in `litellm/__init__.py` (line ~1149) and `litellm/constants.py` (line ~691): `LITELLM_CHAT_PROVIDERS` lists ~70 native chat providers (openai, anthropic, gemini, cohere, together_ai, etc.) plus `openai_compatible_providers` (~20 more like deepseek, groq, perplexity, cerebras) that speak OpenAI's wire format directly. Each provider implements a subclass of `BaseLLM` (`litellm/llms/base.py:14-80`) that overrides `process_response()` and `validate_environment()`.

**Request translation.** Each provider subdirectory under `litellm/llms/` (Anthropic, Bedrock, Gemini, Cohere, etc.) contains a provider module that translates the incoming OpenAI-format request (messages, tools, stream, response_format) into the provider's native format and then maps the response back to `ModelResponse` — LiteLLM's pydantic model that mirrors `openai.types.completion.ChatCompletion`. For example `litellm/llms/anthropic/` handles Anthropic's different tool-call, thinking, and streaming formats.

**Streaming.** Each provider normalizes its stream chunks into OpenAI-format `ChatCompletionChunk` objects via a `CustomStreamWrapper`, so downstream code sees the same SSE event structure. Tool calls and multimodal (image/audio/video inputs mapped to OpenAI content-block arrays) are translated per-provider.

**Cost table.** The 80,644-line `model_prices_and_context_window.json` at the repo root is the authoritative price list with `input_cost_per_token`, `output_cost_per_token`, `max_input_tokens`, `mode`, `supports_*` flags for every known model — parsed at runtime by `cost_calculator.py:1-80`.

**Pass-through routes.** The proxy also supports raw provider-API pass-through for endpoints it hasn't modeled (`litellm/proxy/pass_through_endpoints/`), allowing the proxy to act as a plain credential/rate-limit gate for any provider API.

> **Editor's note.** Correction: providers are implemented as `Config` classes that extend `BaseConfig` (`llms/base_llm/chat/transformation.py`) with `transform_request`/`transform_response`, called by `BaseLLMHTTPHandler`; `BaseLLM.process_response` in `llms/base.py` is a legacy template.

Citations: [litellm/llms/base.py:14-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/llms/base.py#L14-L80) · [litellm/constants.py:691-710](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/constants.py#L691-L710) · [litellm/constants.py:978-992](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/constants.py#L978-L992) · [model_prices_and_context_window.json:1-41](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/model_prices_and_context_window.json#L1-L41) · [litellm/cost_calculator.py:1-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/cost_calculator.py#L1-L80)

### How are API keys, users and tenants managed? (answered)

**Virtual keys (Verification Tokens).** The core auth model is `LiteLLM_VerificationToken` in the Prisma schema (`schema.prisma`), which stores hashed API keys with associated permissions (models, budgets, teams, end-users, metadata). The `user_api_key_auth.py:3487+` function is the FastAPI dependency called on every proxied request — it looks up the bearer token, validates it against the DB (with Redis cache backing via `UserApiKeyCache`), and returns a `UserAPIKeyAuth` object that downstream handlers use for access decisions.

**Multi-tenant hierarchy.** The data model supports a full RBAC hierarchy: **Organizations** → **Teams** → **Users** → **API Keys**, with **Team Memberships** linking Users to Teams with role/permission scoping. Budgets cascade: an Organization budget limits its Teams, a Team budget limits its Members and Keys. The auth check in `_can_object_call_model()` (`litellm/proxy/auth/auth_checks.py`) traverses this chain.

**Upstream credential storage.** Provider API keys are stored in `LiteLLM_CredentialsTable` (`schema.prisma`) and can be resolved from environment variables, AWS Secrets Manager, Google Cloud KMS, HashiCorp Vault, or custom secret managers (`litellm/secret_managers/`).

**Auth methods.** The proxy accepts API keys via the `Authorization: Bearer` header (OpenAI-compatible), `x-api-key` (Anthropic-compatible), or `api-key` (Azure-compatible) — see `litellm/proxy/auth/user_api_key_auth.py:183-199`. It also supports JWT auth (`litellm/proxy/auth/handle_jwt.py`), OAuth2 proxy hooks, and custom SSO.

**Admin UI.** The `/ui` route serves a React-based admin dashboard. Management endpoints under `/key/`, `/user/`, `/team/`, `/organization/` (`litellm/proxy/management_endpoints/`) allow CRUD for all auth entities. A master key (`LITELLM_MASTER_KEY`) bootstraps the first admin.


Citations: [litellm/proxy/auth/user_api_key_auth.py:80-200](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/auth/user_api_key_auth.py#L80-L200) · [litellm/proxy/auth/user_api_key_auth.py:3487-3495](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/auth/user_api_key_auth.py#L3487-L3495) · [litellm/proxy/auth/auth_checks.py:1-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/auth/auth_checks.py#L1-L80) · [litellm/secret_managers/main.py:1-40](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/secret_managers/main.py#L1-L40) · [schema.prisma:1-60](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/schema.prisma#L1-L60)

### How are rate limits, budgets and cost tracking implemented? (answered)

**Rate limits.** The `LiteLLM_BudgetTable` in `schema.prisma` defines `max_parallel_requests`, `tpm_limit` (tokens per minute), `rpm_limit` (requests per minute), and `tpd_limit` (tokens per day) at every level of the hierarchy (org, team, user, key, end-user, tag). These are enforced via the budget throttle middleware (`litellm/proxy/auth/budget_throttle.py`) and pre-call RPM/TPM checks in the routing layer (`router.py:8984-9033`).

**Spend and budget tracking.** `litellm/proxy/spend_tracking/` implements a batched spend counter (`spend_counter_batch.py`) that increments in-memory counters and flushes to PostgreSQL asynchronously on a cadence. The `LiteLLM_SpendLog*` tables in the Prisma schema record every request's token usage, cost, and metadata. Budget enforcement checks (`_virtual_key_max_budget_check`, `_virtual_key_soft_budget_check` in `litellm/proxy/auth/auth_checks.py`) compare accumulated spend against `max_budget`/`soft_budget` thresholds.

**Price table.** `model_prices_and_context_window.json` (80,644 lines) is the bundled pricing catalog. Every model entry specifies `input_cost_per_token`, `output_cost_per_token`, `output_cost_per_reasoning_token`, regional uplift multipliers, prompt-caching discounts, batch rate reductions, and modality-specific costs (audio, image, video, web-search, computer-use). The cost calculator (`litellm/cost_calculator.py`) uses a chain-of-responsibility pattern: each provider module registers a `cost_per_token()` function (e.g. `openai_cost_per_token`, `anthropic_cost_per_token`, `bedrock_cost_per_token` in `litellm/llms/{provider}/cost_calculation.py`), and the base resolver selects the right one.

**Usage accounting.** The final `response_cost`, `total_tokens`, and breakdown are stored in the `LiteLLM_SpendLog*` tables. Budget carry-forward across reset periods is handled by `carried_budget_state.py`. Enterprise users get PTU (pay-per-use) flat-cost pricing via `ptu_flat_cost_rollup.py`.


Citations: [schema.prisma:1-60](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/schema.prisma#L1-L60) · [litellm/router.py:8984-9033](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L8984-L9033) · [litellm/proxy/spend_tracking/spend_counter_batch.py:1-40](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/spend_tracking/spend_counter_batch.py#L1-L40) · [litellm/cost_calculator.py:1-80](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/cost_calculator.py#L1-L80) · [litellm/proxy/auth/auth_checks.py:1-60](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/auth/auth_checks.py#L1-L60)

### How are caching and guardrails implemented? (answered)

**Exact response caching.** The `Cache` class (`litellm/caching/caching.py:74+`) supports multiple backends: in-memory (`InMemoryCache`), Redis (`RedisCache`), Redis Cluster (`RedisClusterCache`), DiskCache, S3, GCS, and Azure Blob. The `DualCache` wrapper maintains a local in-memory LRU + a remote Redis cache for fast local reads with cluster consistency. Response caching is controlled by `CacheMode` (default_on vs default_off) and per-request `ttl`.

**Semantic caching.** Two vector-based semantic caches are available: **QdrantSemanticCache** (`litellm/caching/qdrant_semantic_cache.py`) embeds prompts via a configurable embedding model and queries the Qdrant vector DB for semantically similar requests, returning the cached response if similarity exceeds a threshold. **RedisSemanticCache** (`litellm/caching/redis_semantic_cache.py`) does the same using Redis Stack's vector-similarity-search (VSS) capabilities.

**Guardrails.** The `GuardrailRegistry` in `litellm/proxy/guardrails/guardrail_registry.py:222-386` maintains `guardrail_class_registry` — a dict mapping integration names to `CustomGuardrail` subclasses. ~50 guardrail hooks are auto-discovered from `litellm/proxy/guardrails/guardrail_hooks/` via `get_guardrail_class_from_hooks()` (line 310), which scans subdirectories for `guardrail_class_registry` dicts. Built-in integrations include:
- **PII redaction**: Presidio-based `_OPTIONAL_PresidioPIIMasking` (`guardrail_hooks/presidio.py:172+`) analyzes and anonymizes PII (names, emails, SSNs, credit cards) in request/response content, including SSE stream chunks.
- **Moderation**: Lakera AI (`guardrail_hooks/lakera_ai.py:49`) for prompt injection and content moderation. Also OpenAI Moderation, Bedrock Guardrails, and Guardrails AI.
- **Custom hooks**: Any guardrail can implement `CustomGuardrail` (`litellm/integrations/custom_guardrail.py`) with pre-request (modify/block) and post-response callbacks, plus streaming support for SSE-based anonymization.

**Plugin points.** Guardrails register as standard `litellm.callbacks` via `CustomLogger` hooks, so they also participate in the general callback pipeline (pre-request, post-success, post-failure, streaming).


Citations: [litellm/caching/caching.py:74-100](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/caching/caching.py#L74-L100) · [litellm/caching/qdrant_semantic_cache.py:1-50](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/caching/qdrant_semantic_cache.py#L1-L50) · [litellm/proxy/guardrails/guardrail_registry.py:222-386](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/guardrails/guardrail_registry.py#L222-L386) · [litellm/proxy/guardrails/guardrail_hooks/presidio.py:172-200](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/guardrails/guardrail_hooks/presidio.py#L172-L200)

### How is it observed, deployed and scaled? (answered)

**Observability (OpenTelemetry).** LiteLLM has a comprehensive OTel integration. The `opentelemetry_logger` (`litellm/integrations/otel/logger.py`) implements `CustomLogger` and emits spans for every LLM call, routing decision, guardrail invocation, and service operation. Span attributes include request metadata, deployment info, latency, token counts, and cost. Routing decisions are captured via `routing_decision_attributes()` (`litellm/integrations/otel/routing.py:48-57`) which emits ~40 typed fields per decision. The `tracing/` module (`literllm/tracing/__init__.py`) provides an OTLP HTTP receiver for forwarding spans to ClickHouse or any OTel-compatible backend. Prometheus metrics are exposed at `/metrics` (`litellm/proxy/prometheus_metrics_server.py`).

**Architecture and concurrency.** Pure Python + asyncio. The proxy is a FastAPI app deployed via **uvicorn** (ASGI server, configurable worker count via `--num_workers`, `proxy_cli.py:295-329`). Async at every layer: `httpx.AsyncClient` for outgoing provider calls, `asyncio` for concurrent request handling, and `asyncpg` via Prisma for DB access. Redis caching uses `redis-py` async. Deployment slots use Python `anyio` semaphores for max-parallel-requests enforcement (`router.py:9035-9050`).

**Deployment modes.** The CLI (`litellm/proxy/proxy_cli.py`) uses **click** to launch uvicorn with `--port`, `--host`, `--num_workers`, `--max_requests_before_restart` (for memory leak defense), `--timeout_worker_healthcheck`, and JSON logging support. Official Docker images (`Dockerfile`, `docker-compose.yml`) include PostgreSQL + Redis containers. Helm charts and Terraform scripts are provided for Kubernetes deployments. The `--config` flag points to a YAML/JSON config file defining models, routing, auth, and guardrail settings.

**High availability.** PostgreSQL is the durable backing store (via Prisma ORM with connection pooling via PgBouncer). Redis can be clustered for caching resilience. LiteLLM supports DB-native compaction, query-engine reaping, and graceful worker restarts. Logging is pluggable via `litellm.callbacks` — Langfuse, Datadog, New Relic, and custom loggers are first-class integrations.

> **Editor's note.** Correction: not purely Python. A Rust extension (`litellm-rust/`, loaded as `litellm.rust_bridge._native`) ships in the wheel; per-route rollout rules keep chat and embeddings on Python, let Anthropic messages opt in via `LITELLM_RUST`, and require Rust for OCR and Bedrock transcription.

Citations: [litellm/integrations/otel/logger.py:1-60](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/integrations/otel/logger.py#L1-L60) · [litellm/integrations/otel/routing.py:48-57](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/integrations/otel/routing.py#L48-L57) · [litellm/proxy/proxy_cli.py:295-329](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/proxy_cli.py#L295-L329) · [litellm/router.py:9035-9050](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L9035-L9050)
