# How is it observed, deployed and scaled?

> LLM gateways — a good answer covers: Logs, metrics and traces (OpenTelemetry); performance architecture (language, concurrency); deployment modes; HA.

Canonical page: https://llms-technical-reviews.com/llm-gateways/q/ops/

## Verdict

[Agent Router](/p/agent-router/) and [Higress](/p/higress/) scale most like ordinary infrastructure, because Envoy carries the traffic. [Bifrost](/p/bifrost/) is the most complete single binary for metrics and tracing. This site runs no benchmarks. The speed figures below are the projects' own claims, and only the architecture is compared.

**Envoy data planes.** Agent Router injects a Go ext_proc into each Envoy pod, as a regular container or as a native sidecar. It exports OpenTelemetry metrics (Prometheus always on) and traces with GenAI conventions. Each request makes gRPC round trips to that sidecar. Higress runs Go plugins compiled to WASM inside Envoy, with an Istio-based controller. `ai-statistics` emits Envoy counters per route, model and consumer. [Plano](/p/plano/) pairs Envoy and Rust WASM filters with a separate async Rust service, so a request crosses Envoy twice. It exports OTLP traces, and Redis lets replicas share session state.

**Go services.** Bifrost gives each provider a bounded queue with a fixed worker pool and uses fasthttp and object pools. It has Prometheus and OpenTelemetry plugins. Its README reports 11 µs of added overhead at 5,000 RPS. Counters and session affinity are per node in the open-source build. [One API](/p/one-api/) and [New API](/p/new-api/) are Gin binaries on SQLite, MySQL or PostgreSQL. Replicas set `NODE_TYPE=slave` and share the database and Redis. Neither exports OpenTelemetry, and New API can send logs to ClickHouse. [GPT-Load](/p/gpt-load/) swaps its config snapshot atomically. Its RPM limits, concurrency limits and health state stay in one process.

**Python and Node services.** [LiteLLM](/p/litellm/) runs FastAPI under uvicorn workers. Its chat path is still Python, and only a few routes use its Rust extension. Shared limits need PostgreSQL and Redis. It has Prometheus, OpenTelemetry and many logging callbacks, and its README claims 8 ms P95 latency at 1,000 RPS. [Portkey Gateway](/p/portkey-gateway/) is a stateless Hono app for Node or Cloudflare Workers. Its README claims under 1 ms of latency. Its logs go to `console` and an SSE stream, with no exporter. [OmniRoute](/p/omniroute/) is a Next.js app on SQLite. Its OTLP exporter is off unless an endpoint is set, and its in-memory rotation stores are not shared between replicas.

Pick: Agent Router or Higress on Kubernetes with an Envoy fleet.
Pick: Bifrost for a single binary with built-in Prometheus and OpenTelemetry.
Pick: LiteLLM when logging integrations matter more than per-instance overhead.

## Per-project answers

### diegosouzapw/OmniRoute (answered)

Observability is three-tiered. Structured JSON logging uses pino through `src/lib/logger/`, with a log-export system (`src/lib/logExport/`) that can ship call logs to BigQuery or other destinations. OpenTelemetry support is optional and lightweight: `open-sse/services/routing/otel.ts` exports GenAI semantic-convention spans (gen_ai.provider.name, gen_ai.request.model, gen_ai.usage.*) to an OTLP/HTTP collector. It uses global `fetch` (no SDK dependency), a bounded buffer with background flush, and is disabled by default unless `OMNIROUTE_OTEL_ENDPOINT` or `OTEL_EXPORTER_OTLP_ENDPOINT` is set. Internal metrics (latency, request counts, error rates) are collected in `src/lib/logger/metrics.ts`. Deployment supports multiple modes: container via `Dockerfile` (multi-stage Node.js), Fly.io via `fly.toml`, Docker Compose, and tunnel-based deployment (`src/mitm/tunnel/`) for local model endpoints. The dashboard at `src/app/(dashboard)/` provides visual management of providers, API keys, usage, logs, and settings. Performance architecture is standard Node.js async I/O on Next.js 16, single-threaded event loop with 161 executor modules and connection pooling via Undici (`proxyDispatcherCache.ts` with round-robin dispatchers to prevent SSE stream monopolization). HA mechanisms include the three resilience layers: provider circuit breakers (persisted to SQLite for restart survival), connection cooldown, and model lockout — all lazily self-recovering. Rate limiting uses Redis when available (falling back to in-memory). The `_artifacts/` directory stores temporary working files per CLAUDE.md conventions. TypeScript strict mode is off (target ES2022), with 193 DB migrations for schema evolution.


Citations: [open-sse/services/routing/otel.ts:1-50](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/services/routing/otel.ts#L1-L50) · [open-sse/utils/proxyDispatcherCache.ts:26-50](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/open-sse/utils/proxyDispatcherCache.ts#L26-L50) · [src/shared/utils/circuitBreaker.ts:1-20](https://github.com/diegosouzapw/OmniRoute/blob/8ad6b1c46eaea49ab6b6e9929817c08a90c5067b/src/shared/utils/circuitBreaker.ts#L1-L20)

### BerriAI/litellm (answered)

**Observability (OpenTelemetry).** LiteLLM has a comprehensive OTel integration. The `opentelemetry_logger` (`litellm/integrations/otel/logger.py`) implements `CustomLogger` and emits spans for every LLM call, routing decision, guardrail invocation, and service operation. Span attributes include request metadata, deployment info, latency, token counts, and cost. Routing decisions are captured via `routing_decision_attributes()` (`litellm/integrations/otel/routing.py:48-57`) which emits ~40 typed fields per decision. The `tracing/` module (`literllm/tracing/__init__.py`) provides an OTLP HTTP receiver for forwarding spans to ClickHouse or any OTel-compatible backend. Prometheus metrics are exposed at `/metrics` (`litellm/proxy/prometheus_metrics_server.py`).

**Architecture and concurrency.** Pure Python + asyncio. The proxy is a FastAPI app deployed via **uvicorn** (ASGI server, configurable worker count via `--num_workers`, `proxy_cli.py:295-329`). Async at every layer: `httpx.AsyncClient` for outgoing provider calls, `asyncio` for concurrent request handling, and `asyncpg` via Prisma for DB access. Redis caching uses `redis-py` async. Deployment slots use Python `anyio` semaphores for max-parallel-requests enforcement (`router.py:9035-9050`).

**Deployment modes.** The CLI (`litellm/proxy/proxy_cli.py`) uses **click** to launch uvicorn with `--port`, `--host`, `--num_workers`, `--max_requests_before_restart` (for memory leak defense), `--timeout_worker_healthcheck`, and JSON logging support. Official Docker images (`Dockerfile`, `docker-compose.yml`) include PostgreSQL + Redis containers. Helm charts and Terraform scripts are provided for Kubernetes deployments. The `--config` flag points to a YAML/JSON config file defining models, routing, auth, and guardrail settings.

**High availability.** PostgreSQL is the durable backing store (via Prisma ORM with connection pooling via PgBouncer). Redis can be clustered for caching resilience. LiteLLM supports DB-native compaction, query-engine reaping, and graceful worker restarts. Logging is pluggable via `litellm.callbacks` — Langfuse, Datadog, New Relic, and custom loggers are first-class integrations.

> **Editor's note.** Correction: not purely Python. A Rust extension (`litellm-rust/`, loaded as `litellm.rust_bridge._native`) ships in the wheel; per-route rollout rules keep chat and embeddings on Python, let Anthropic messages opt in via `LITELLM_RUST`, and require Rust for OCR and Bedrock transcription.

Citations: [litellm/integrations/otel/logger.py:1-60](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/integrations/otel/logger.py#L1-L60) · [litellm/integrations/otel/routing.py:48-57](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/integrations/otel/routing.py#L48-L57) · [litellm/proxy/proxy_cli.py:295-329](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/proxy/proxy_cli.py#L295-L329) · [litellm/router.py:9035-9050](https://github.com/BerriAI/litellm/blob/62dee3d73046fb717f8693a7660d2fe457c9370f/litellm/router.py#L9035-L9050)

### QuantumNous/new-api (answered)

**Observability.** The project has **no OpenTelemetry integration** — there are no trace exporters, metric collectors, or OTel SDK references in the code. Observability is instead built on: (1) **File-based logging** (`logger/logger.go`) — `SetupLogger()` writes to a timestamped log file in the configured `--log-dir`, rotating at startup. The `LogInfo`/`LogWarn`/`LogError`/`LogDebug` functions write structured messages with context from the Gin request context. (2) **Request logging** (`model/log.go`) — every API request is recorded in the `Log` table (or ClickHouse log DB), storing tokens, quota, channel, latency (`UseTime`), IP, model, and extended billing metadata. This serves as the audit trail and usage analytics source. (3) **Performance metrics** (`pkg/perf_metrics/`) — relay latency and business rejection rates are captured at the controller level (`controller/relay.go:134`). (4) **Active connection stats** (`middleware/stats.go`) — an atomic counter tracks concurrent requests. (5) **Admin dashboard** — the React UI provides real-time usage dashboards, log exploration, and performance charts.

**Performance architecture.** The project is written in **Go** (compiled, with a green-thread concurrency model via `gopool`). It uses the **Gin** web framework, which is an async Go framework. The thread pool handles concurrent HTTP relay connections efficiently — each relay request is a lightweight goroutine. **Redis** offloads rate limiting state, channel polling indices, and cache lookups, reducing database pressure. The database layer supports **ClickHouse** as a separate log database (`docker-compose.yml:32`) to isolate analytics writes from transactional data.

**Deployment modes.** (1) **Docker** — the primary deployment path (`Dockerfile`) uses a multi-stage build: Bun builds the React frontend, Go compiles the binary, and a Debian slim image runs it on port 3000. The `docker-compose.yml` bundles the app with PostgreSQL (or MySQL) and Redis, with config via environment variables. (2) **Standalone binary** — `go build` produces a single static binary with embedded web frontend. (3) **Electron desktop wrapper** — `electron/` provides a desktop app wrapping the same web UI. (4) **systemd service** (`new-api.service`) is included for Linux native deployment.

**High availability.** The architecture is **single-node** — there is no built-in clustering or HA mode. However, the system is **stateless** (session state in Redis, config in database), so multiple instances can share the same Postgres+Redis backend. The `NODE_NAME` env var (`docker-compose.yml:38`) identifies nodes in audit logs. The `SESSION_SECRET` must be shared across nodes (`docker-compose.yml:41`). Graceful connection handling is configurable via `STREAMING_TIMEOUT` and `RELAY_IDLE_CONN_TIMEOUT`.

**Deployment configuration** is entirely environment-variable driven: `SQL_DSN`, `LOG_SQL_DSN`, `REDIS_CONN_STRING`, `TZ`, `BATCH_UPDATE_ENABLED`, `ERROR_LOG_ENABLED`, plus session security settings (`SESSION_SECRET`, `SESSION_COOKIE_SECURE`, `SESSION_COOKIE_TRUSTED_URL`).


Citations: [logger/logger.go:1-74](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/logger/logger.go#L1-L74) · [Dockerfile:1-40](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/Dockerfile#L1-L40) · [docker-compose.yml:1-50](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/docker-compose.yml#L1-L50) · [controller/relay.go:127-139](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/controller/relay.go#L127-L139) · [model/log.go:59-80](https://github.com/QuantumNous/new-api/blob/973cf8ef4600947a4270e95ada7916740fa8264c/model/log.go#L59-L80)

### songquanpeng/one-api (answered)

**Language and concurrency.** One API is written in Go using the Gin web framework (Go's analog of Express/Flask). Go's goroutine-based concurrency model and the Gin framework's non-blocking I/O provide good per-instance throughput. The server starts in `main.go:103-120`: a single Gin engine with middleware stack (session store, request ID, language, logger, rate limits) and route registration. No explicit worker count or multi-process configuration is visible — Gin runs as a single process with goroutine-based concurrency.

**Logging.** The `middleware/logger.go` sets up standard Gin request logging (method, path, status code, latency, client IP). Application-level logging uses `common/logger/logger.go` with log levels (Info, Error, Debug, SysLog). `config.DebugEnabled` and `config.DebugSQLEnabled` toggle verbose logging and SQL query logging. `ONLY_ONE_LOG_FILE` controls log file rotation (`config.go:159`).

**Metrics and observability.** No OpenTelemetry, Prometheus, or Datadog integration is present. The only built-in metric is the channel success-rate monitor (`monitor/metric.go`), which uses in-memory ring buffers per channel. There are no external metrics exporters.

**Database.** The system supports SQLite (default, file-based), MySQL, and PostgreSQL (`model/main.go:67-109`). Schema is auto-migrated via GORM's `AutoMigrate` on startup. A separate `LOG_SQL_DSN` can route usage logs to a secondary database. Connection pooling is configured via `SQL_MAX_IDLE_CONNS`, `SQL_MAX_OPEN_CONNS`, and `SQL_MAX_LIFETIME` environment variables (`model/main.go:214-217`).

**Deployment.** A multi-stage Dockerfile builds static Go binary + React frontends. The `docker-compose.yml` provides a full stack with Redis, MySQL, and the One API service. Environment variables control all configuration: `SQL_DSN`, `REDIS_CONN_STRING`, `SESSION_SECRET`, `NODE_TYPE`, `SYNC_FREQUENCY`, etc.

**High-availability / multi-node.** The `NODE_TYPE` environment variable (`config.go:105`) supports "master" vs "slave" nodes. Only the master node runs database migrations. Slave nodes can serve API traffic but delegate configuration reads to the database and rely on `SYNC_FREQUENCY` to refresh their in-memory channel cache. The `FRONTEND_BASE_URL` env var allows slave nodes to offload the React frontend to a separate server. There is no load balancer or clustering logic in the application itself — HA relies on an external LB + shared DB/Redis.

**Health checks.** The docker-compose defines a healthcheck at `http://localhost:3000/api/status`. The `/api/status` endpoint (`controller/misc.go`) provides a simple status probe.


Citations: [main.go:29-124](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/main.go#L29-L124) · [model/main.go:67-109](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/model/main.go#L67-L109) · [common/config/config.go:105-111](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/common/config/config.go#L105-L111) · [monitor/metric.go:1-79](https://github.com/songquanpeng/one-api/blob/8df4a2670b98266bd287c698243fff327d9748cf/monitor/metric.go#L1-L79)

### Portkey-AI/gateway (answered)

**Logging** is in-memory and SSE-streaming. The `logHandler` middleware (`src/middlewares/log/index.ts:149`) captures request/response data into a `requestOptions` array on the Hono context. After each request, `processLog` (line 112) sanitizes sensitive headers and broadcasts the log entry to all connected SSE clients at `/log/stream`. The log stream is protected by `adminAuthMiddleware` and includes timing, provider, status, and truncated response bodies (capped at 100KB, line 6).

**Metrics and traces** are minimal. The `LogsService` at `src/handlers/services/logsService.ts:96-162` can construct OTLP-format span objects for tool execution events (with `traceId`, `spanId`, `parentSpanId`, and `gen_ai.*` attributes), but these are not exported to any OTel collector by the gateway — the structure is defined but the exporter is not implemented. The logger is plain `console` (`src/apm/index.ts:1`).

**Performance architecture**: Built on Hono (ultra-fast web framework designed for edge), written in TypeScript. **Concurrency** is handled by Hono's async middleware pipeline and Cloudflare Workers' `waitUntil` for fire-and-forget logging. There is no per-request thread pool — on Node.js it uses the async event loop, and on Workerd it uses the Workers runtime. Hono's sub-millisecond routing overhead is a key design goal (README claim, but the architecture supports it).

**Deployment modes**: Three runtimes are supported: **Cloudflare Workers** (`npm run dev` on wrangler), **Node.js** (`npm run dev:node` or `npm run start:node`), and any Hono-compatible runtime (Bun, Deno via adapter). The `getRuntimeKey()` check in `src/index.ts` and `src/middlewares/log/index.ts` switches behavior per runtime. The Node.js server (`src/start-server.ts`) uses `@hono/node-server` and supports WebSocket upgrades for `/v1/realtime`, SSE log streaming, and a local admin UI at `/public/` that can be hidden with `--headless`. A `conf.json` file controls plugins, caching, and admin credentials.

**High Availability**: The open-source gateway itself has no built-in HA — deployments rely on the underlying platform (Cloudflare Workers' global network, Node.js process manager, Docker orchestrator). Redis provides a shared cache for multi-instance deployments.


Citations: [src/middlewares/log/index.ts:80-166](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/middlewares/log/index.ts#L80-L166) · [src/handlers/services/logsService.ts:96-162](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/handlers/services/logsService.ts#L96-L162) · [src/index.ts:47-51](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/index.ts#L47-L51) · [src/start-server.ts:1-40](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/start-server.ts#L1-L40) · [src/apm/index.ts:1-1](https://github.com/Portkey-AI/gateway/blob/669825cbe89ee51569918b8f78a9db486fd69dd4/src/apm/index.ts#L1-L1)

### higress-group/higress (answered)

**Higress provides extensive observability, deploys as a Docker container or via Helm on Kubernetes, and runs as a high-performance Envoy-based proxy.**

**Logs and traces**: The `ai-statistics` plugin writes structured AI logs (`ai_log`) and OpenTelemetry-style trace spans. At the request header phase, it sets the span kind `gen_ai.span.kind = LLM` and extracts the route, cluster, API name — `ai-statistics/main.go:700-750`. On the response, it records model name, input/output/total tokens, `llm_first_token_duration`, and `llm_service_duration` — `ai-statistics/main.go:1069-1098`. Span attributes use the `gen_ai.*` prefix for ARMS (Alibaba Resource Monitoring System) compatibility — `ai-statistics/main.go:102-108`. Built-in attributes like `question`, `answer`, `reasoning`, `tool_calls`, `system`, `reasoning_tokens`, `cached_tokens`, and full `input_token_details`/`output_token_details` maps are extracted from both OpenAI and Anthropic response formats — `ai-statistics/main.go:115-144`. Streaming responses are re-assembled using an SSE framer for accurate token extraction — `ai-statistics/main.go:895-919`.

**Metrics**: The plugin defines Envoy-native counter metrics at the granularity of `route.{route}.upstream.{cluster}.model.{model}.consumer.{consumer}.metric.{metricName}` — `ai-statistics/main.go:481-483`. Metrics include token counts, request duration, stream duration count, and failure count. Error detection works across both streaming and non-streaming responses, checking response-level error fields and HTTP status codes — `ai-statistics/main.go:1519-1560`.

**Performance architecture**: Higress is built on Envoy (C++), with extension logic running as **WASM plugins** compiled from Go using TinyGo — `plugins/wasm-go/Makefile:1-8`. The WASM sandbox provides process-level isolation. Request body buffering is configurable per-extension (default 100MB) — `ai-proxy/main.go:30`. The WASM VM memory limit triggers automatic VM rebuild when exceeded — `ai-endpoint-picker/main.go:230-238`. Lua scripts in Redis are used for atomic rate-limit operations — `ai-token-ratelimit/main.go:58-79`.

**Deployment modes**: The quickstart uses Docker with included configuration (`docker run -d --rm --name higress-ai`) — `README.md:70-72`. Kubernetes deployment uses Helm charts and CustomResourceDefinitions for plugin configuration management — `helm/`. The API proto definitions at `api/networking/v1/` define the control plane data models — `api/networking/v1/http_2_rpc.proto`.

**HA and scaling**: The failover mechanism (`failover.go`) uses VM-level leader election (CAS-based lease) to coordinate health checks across Wasm VMs — `ai-proxy/provider/failover.go:292-341`. Token failover tracks per-token health with configurable failure/success thresholds and cooldown recovery — `ai-proxy/provider/failover.go:349-403`. The `x-higress-fallback-from` header enables safe internal redirects across the gateway chain — `ai-proxy/main.go:189-196`.


Citations: [plugins/wasm-go/extensions/ai-statistics/main.go:700-750](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L700-L750) · [plugins/wasm-go/extensions/ai-statistics/main.go:1069-1098](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L1069-L1098) · [plugins/wasm-go/extensions/ai-statistics/main.go:481-524](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-statistics/main.go#L481-L524) · [plugins/wasm-go/extensions/ai-proxy/provider/failover.go:292-341](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/failover.go#L292-L341) · [plugins/wasm-go/extensions/ai-proxy/provider/failover.go:349-403](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/provider/failover.go#L349-L403) · [plugins/wasm-go/extensions/ai-proxy/main.go:189-196](https://github.com/higress-group/higress/blob/bda81f1067e1285775e650644a20a635f51b0a6d/plugins/wasm-go/extensions/ai-proxy/main.go#L189-L196)

### maximhq/bifrost (answered)

**Observability:** Bifrost has three observability layers. (1) Prometheus metrics via the `telemetry` plugin (`plugins/telemetry/main.go`) — exposes counters for upstream requests, success/error rates, input/output tokens, cache hits, and cost; histograms for upstream latency, overhead latency, stream first-token/inter-token latency, and retry counts. Supports push to a Prometheus Push Gateway for multi-node aggregation. (2) OpenTelemetry via the `otel` plugin (`plugins/otel/main.go`) — exports traces and metrics to one or more collector profiles (OTLP over HTTP or gRPC). Each profile has independent endpoints, headers, TLS config, and span filters. Traces follow the GenAI semantic convention and include per-provider spans, fallback spans, and overhead breakdowns. Supports distributed tracing via W3C traceparent and session-based trace grouping. (3) Structured logging via the `logging` plugin (`plugins/logging/main.go`) — logs all requests/responses to PostgreSQL with search, filter, and pagination via the `logstore` framework. **Performance architecture:** Written in Go with fasthttp (a high-performance async HTTP client/server), object pooling for responses, lock-free sync.Map data structures in hot paths, and per-provider connection pools with configurable concurrency. The model catalog uses memoization with generation counters to avoid recomputation. **Deployment:** Docker images (Alpine-based multi-stage builds embedding a React UI) run the HTTP server on port 8080 by default. Helm charts (`helm-charts/bifrost/`) deploy to Kubernetes with configurable replica count, secret management, and values examples for various backing stores (PostgreSQL+Weaviate, SQLite+Redis, etc.). The Dockerfile (`transports/Dockerfile`) builds a static Go binary using CGO. **High availability:** Horizontal scaling is supported by shared PostgreSQL and vector stores; the KV store (`framework/kvstore/`) provides distributed state for session affinity and rate-limit coordination across nodes. The OTel circuit breaker (`plugins/otel/main.go:56-64`) prevents a failed collector from blocking request processing.

> **Editor's note.** Correction: framework/kvstore is an in-memory, per-process store (kvstore.go L25-L33), and governance usage counters are also kept per process, so session affinity and rate limits are not coordinated across nodes in the open-source build; the README lists clustering as an enterprise feature.

Citations: [plugins/telemetry/main.go:1-50](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/telemetry/main.go#L1-L50) · [plugins/otel/main.go:1-70](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/otel/main.go#L1-L70) · [plugins/otel/metrics.go:36-80](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/otel/metrics.go#L36-L80) · [plugins/logging/main.go:1-35](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/plugins/logging/main.go#L1-L35) · [transports/Dockerfile:1-40](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/transports/Dockerfile#L1-L40) · [helm-charts/bifrost/values.yaml:1-50](https://github.com/maximhq/bifrost/blob/0e9c135bc16e49aabaafc58aaea0a6777a836634/helm-charts/bifrost/values.yaml#L1-L50)

### katanemo/plano (answered)

**Plano is observed through OpenTelemetry tracing and structured logging, deployed as a Docker container with supervisord managing three processes, and scales horizontally behind a load balancer.**

**Logs, metrics, and traces**: OpenTelemetry is the primary observability pillar. The `tracing` module in `crates/brightstaff/src/tracing/` initializes a tracer (`init.rs`) with configurable sampling rate and OpenTelemetry gRPC endpoint (`configuration.rs:480-485`). Custom span attributes are injected from HTTP headers via `collect_custom_trace_attributes()`. LLM-specific semantic conventions are defined in `tracing/constants.rs:49` (model name, provider, request/response content length, temperature, token counts). A dedicated `PostHogExporter` (`tracing/posthog_exporter.rs:43`) translates LLM spans into `$ai_generation` events and POSTs them to PostHog's `batch/` API, supporting `distinct_id` from a configurable request header. The `otlp_exporter` gRPC target is also natively supported via Envoy's OpenTelemetry tracing integration (`envoy.template.yaml:47-57`).

**Metrics** are instrumented at two levels: WASM-level counters/gauges via `proxy_wasm::hostcalls` (`crates/common/src/stats.rs`) and application-level metrics collected throughout brightstaff (`crates/brightstaff/src/metrics/`). The `time_to_first_token` histogram is configured in `envoy.template.yaml:7-28`. Access logs are written to `/var/log/access_ingress.log` in standard Envoy format.

**Performance architecture**: The system is written in **Rust** (edition 2021, safe Rust throughout). The two WASM plugins (`prompt_gateway`, `llm_gateway`) are `cdylib` crates compiled for `wasm32-wasip1` and run inside Envoy's WASM sandbox — meaning synchronous, callback-driven execution with no tokio/async, no networking, and no filesystem access. The native `brightstaff` binary runs as a separate async Tokio process for routing, state, and tracing. Communication between Envoy and brightstaff happens over HTTP; brightstaff exposes an orchestrator API (`orchestrator_url`) for routing decisions.

**Deployment**: A single Docker image bundles all components. `supervisord.conf` orchestrates three programs: `config_generator` (Python CLI, runs once to render YAML configs), `brightstaff` (native binary), and `envoy` (the proxy). The config generator runs `envsubst` on rendered templates to inject environment variables. The `planoai` Python CLI (`cli/planoai/main.py`) provides `up`, `down`, `build`, `logs`, `trace` and `init` commands for local development and deployment management.

**High availability**: Session bindings can be stored in Redis (`crates/brightstaff/src/session_cache/redis.rs`) rather than in-memory, enabling multiple brightstaff instances to share session state behind a load balancer. The Redis `ConnectionManager` with automatic reconnection handles connection drops. Session-cache lookups use short timeouts (`connect: 5s, response: 2s`) so a slow backend never blocks routing decisions. The project does not implement active health checks — it relies on Envoy's upstream cluster health monitoring and the response code fallback chain. Key rotation is handled through env-substituted configs and `passthrough_auth` mode.


Citations: [crates/brightstaff/src/tracing/init.rs:1-15](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/tracing/init.rs#L1-L15) · [crates/brightstaff/src/tracing/posthog_exporter.rs:1-51](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/tracing/posthog_exporter.rs#L1-L51) · [crates/brightstaff/src/tracing/constants.rs:1-61](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/tracing/constants.rs#L1-L61) · [crates/brightstaff/src/session_cache/redis.rs:1-35](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/brightstaff/src/session_cache/redis.rs#L1-L35) · [crates/common/src/configuration.rs:479-528](https://github.com/katanemo/plano/blob/72002a62d90ad13dd246cf3ef8d99d8c98b075ff/crates/common/src/configuration.rs#L479-L528)

### tbphp/gpt-load (answered)

**Logging** — the system uses `logrus` with structured fields. Request logs are written via `telemetry.RequestLogSink` (`handler.go:485-510`) — every request produces a `requestRecorder` that collects operation, model, cost, status, and timing. The `requestlog` package (`internal/requestlog/`) persists logs with configurable retention (default 7 days, `maxRequestLogRetentionDays` 365, `internal/state/runtime_settings.go:59-61`). Multiple log stores are supported: access-key usage, credential window usage, group usage, RPM history, and quota history.

**Metrics** — `health.StatsStore` (`internal/health/stats.go:14`) tracks per-credential success/failure/problem counts in sliding 5-minute windows (1-minute buckets). RPM counters are stored in `rpm.Store` and queried via `ratelimit.AccessKeyRPM`. No OpenTelemetry export is implemented — metrics are consumed in-process for health decisions and exposed via the control API.

**Traces** — no distributed tracing (OpenTelemetry or similar). The system uses request IDs (`handler.newRequestID()`, `handler.go:464`) that flow in `X-GPTLoad-Request-ID` response headers and appear in all log entries, enabling request correlation.

**Deployment** — packaged as a single Go binary via Docker (`Dockerfile`). The build uses multi-stage: Node.js for the web frontend, then Go compilation. The final image is Alpine-based. A `docker-compose.yml` and `docker-compose.voice.yml` are provided for container orchestration. The binary supports Windows as a first-class target with service management (`service_windows.go`, `service_cli.go`).

**Concurrency model** — Go language with goroutines (green threads). The `gin` HTTP framework provides non-blocking request handling. The `concurrency.go` package tracks active requests per access-key and globally. `channel.RouteMode` and the scheduler are pure in-memory operations with no per-request blocking IO beyond the upstream HTTP call.

**Configuration management** — `state.Manager` (`internal/state/manager.go:11`) holds an atomic pointer to the current `ConfigSnapshot`. Snapshots are immutable and replaced atomically on publish via `sync/atomic.Pointer`. The `SnapshotReconciler` interface synchronizes infrastructure resources before a snapshot becomes visible. Configuration is loaded from a database via `state/loader` and refreshed in-place without restarts.

**High availability** — the system is single-process; there is no built-in clustering, leader election, or shared state between instances. HA would rely on running multiple instances behind a load balancer with a shared database for configuration persistence. RPM and concurrency limits are in-process and not shared across nodes.

**Health endpoints** — the `health` package exposes credential health stats via the control API, and the auth layer checks credential cooldown/blacklist state inline during scheduling.


Citations: [internal/health/stats.go:1-60](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/health/stats.go#L1-L60) · [internal/gateway/handler.go:94-112](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/gateway/handler.go#L94-L112) · [internal/state/manager.go:1-60](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/internal/state/manager.go#L1-L60) · [Dockerfile:1-30](https://github.com/tbphp/gpt-load/blob/a5c691bb0559f856ebb587879d82e341ff480202/Dockerfile#L1-L30)

### theagentrouter/agent-router (answered)

The project is observed via OpenTelemetry for both metrics and traces. **Metrics** are backed by OpenTelemetry with configurable exporters: Prometheus (always enabled), console exporter, and OTLP exporter. Environment variables `OTEL_METRICS_EXPORTER`, `OTEL_EXPORTER_OTLP_ENDPOINT`, and `OTEL_SDK_DISABLED` control configuration (`internal/metrics/metrics.go:35-94`). The `Metrics` interface tracks request start/completion timing, token usage, backend selection, time-to-first-token, and inter-token latency for streaming (`internal/metrics/metrics.go:97-127`). **Traces** support the same exporters (console, OTLP) with `OTEL_TRACES_EXPORTER` and auto-propagation via `autoprop`. Span recorders exist for every endpoint type (chat completions, embeddings, responses, speech, transcription, MCP, etc.) with semantic conventions for both OpenInference and OTEL GenAI (`internal/tracing/tracing.go:137-328`). **Performance architecture**: The Go-based ext_proc sidecar handles request/response transformations per-stream via gRPC. Router processors are tracked per-request-ID with a read-write mutex for concurrent access (`internal/extproc/server.go:54-64`). The data plane itself runs on Envoy Proxy (C++), providing battle-tested HTTP/2, connection pooling, and async I/O. **Deployment modes**: (1) Standalone CLI: `aigw run` downloads and configures an Envoy binary with the ext_proc sidecar as a local process (`cmd/aigw/main.go:38-40`). (2) Kubernetes: runs as a controller-runtime operator that reconciles Gateway resources, injects the ext_proc as a sidecar container (DaemonSet or init container), and manages rolling updates via workload template annotations (`internal/controller/gateway.go:1298-1391`). (3) Admin server: `AdminPort` exposes `/metrics` and `/health` HTTP endpoints (`cmd/aigw/main.go:49`). **HA**: Kubernetes controller pattern provides self-healing via reconciliation; Envoy's data plane handles connection retries and failover. Config bundle checksums ensure data-plane integrity (`internal/filterapi/config_bundle.go:21-91`).

> **Editor's note.** Correction: the ext_proc is injected into Envoy pods either as a regular container or, in sidecar mode, as a native sidecar (init container with restartPolicy: Always); it is not deployed as a DaemonSet.

Citations: [internal/metrics/metrics.go:35-94](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/metrics/metrics.go#L35-L94) · [internal/metrics/metrics.go:97-127](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/metrics/metrics.go#L97-L127) · [internal/tracing/tracing.go:137-328](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/tracing/tracing.go#L137-L328) · [internal/extproc/server.go:54-64](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/extproc/server.go#L54-L64) · [internal/controller/gateway.go:1298-1391](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/internal/controller/gateway.go#L1298-L1391) · [cmd/aigw/main.go:45-51](https://github.com/theagentrouter/agent-router/blob/daa9f891a8afcb18576d4593ad18870d1d18abac/cmd/aigw/main.go#L45-L51)
