How is it observed, deployed and scaled?
Logs, metrics and traces (OpenTelemetry); performance architecture (language, concurrency); deployment modes; HA.
Verdict
Agent Router and Higress scale most like ordinary infrastructure, because Envoy carries the traffic. Bifrost is the most complete single binary for metrics and tracing. This site runs no benchmarks. The speed figures below are the projects’ own claims, and only the architecture is compared.
Envoy data planes. Agent Router injects a Go ext_proc into each Envoy pod, as a regular container or as a native sidecar. It exports OpenTelemetry metrics (Prometheus always on) and traces with GenAI conventions. Each request makes gRPC round trips to that sidecar. Higress runs Go plugins compiled to WASM inside Envoy, with an Istio-based controller. ai-statistics emits Envoy counters per route, model and consumer. Plano pairs Envoy and Rust WASM filters with a separate async Rust service, so a request crosses Envoy twice. It exports OTLP traces, and Redis lets replicas share session state.
Go services. Bifrost gives each provider a bounded queue with a fixed worker pool and uses fasthttp and object pools. It has Prometheus and OpenTelemetry plugins. Its README reports 11 µs of added overhead at 5,000 RPS. Counters and session affinity are per node in the open-source build. One API and New API are Gin binaries on SQLite, MySQL or PostgreSQL. Replicas set NODE_TYPE=slave and share the database and Redis. Neither exports OpenTelemetry, and New API can send logs to ClickHouse. GPT-Load swaps its config snapshot atomically. Its RPM limits, concurrency limits and health state stay in one process.
Python and Node services. LiteLLM runs FastAPI under uvicorn workers. Its chat path is still Python, and only a few routes use its Rust extension. Shared limits need PostgreSQL and Redis. It has Prometheus, OpenTelemetry and many logging callbacks, and its README claims 8 ms P95 latency at 1,000 RPS. Portkey Gateway is a stateless Hono app for Node or Cloudflare Workers. Its README claims under 1 ms of latency. Its logs go to console and an SSE stream, with no exporter. OmniRoute is a Next.js app on SQLite. Its OTLP exporter is off unless an endpoint is set, and its in-memory rotation stores are not shared between replicas.
Pick: Agent Router or Higress on Kubernetes with an Envoy fleet. Pick: Bifrost for a single binary with built-in Prometheus and OpenTelemetry. Pick: LiteLLM when logging integrations matter more than per-instance overhead.
Per-project answers
diegosouzapw/OmniRoute
answeredObservability is three-tiered. Structured JSON logging uses pino through src/lib/logger/, with a log-export system (src/lib/logExport/) that can ship call logs to BigQuery or other destinations. OpenTelemetry support is optional and lightweight: open-sse/services/routing/otel.ts exports GenAI semantic-convention spans (gen_ai.provider.name, gen_ai.request.model, gen_ai.usage.*) to an OTLP/HTTP collector. It uses global fetch (no SDK dependency), a bounded buffer with background flush, and is disabled by default unless OMNIROUTE_OTEL_ENDPOINT or OTEL_EXPORTER_OTLP_ENDPOINT is set. Internal metrics (latency, request counts, error rates) are collected in src/lib/logger/metrics.ts. Deployment supports multiple modes: container via Dockerfile (multi-stage Node.js), Fly.io via fly.toml, Docker Compose, and tunnel-based deployment (src/mitm/tunnel/) for local model endpoints. The dashboard at src/app/(dashboard)/ provides visual management of providers, API keys, usage, logs, and settings. Performance architecture is standard Node.js async I/O on Next.js 16, single-threaded event loop with 161 executor modules and connection pooling via Undici (proxyDispatcherCache.ts with round-robin dispatchers to prevent SSE stream monopolization). HA mechanisms include the three resilience layers: provider circuit breakers (persisted to SQLite for restart survival), connection cooldown, and model lockout — all lazily self-recovering. Rate limiting uses Redis when available (falling back to in-memory). The _artifacts/ directory stores temporary working files per CLAUDE.md conventions. TypeScript strict mode is off (target ES2022), with 193 DB migrations for schema evolution.
BerriAI/litellm
answeredObservability (OpenTelemetry). LiteLLM has a comprehensive OTel integration. The opentelemetry_logger (litellm/integrations/otel/logger.py) implements CustomLogger and emits spans for every LLM call, routing decision, guardrail invocation, and service operation. Span attributes include request metadata, deployment info, latency, token counts, and cost. Routing decisions are captured via routing_decision_attributes() (litellm/integrations/otel/routing.py:48-57) which emits ~40 typed fields per decision. The tracing/ module (literllm/tracing/__init__.py) provides an OTLP HTTP receiver for forwarding spans to ClickHouse or any OTel-compatible backend. Prometheus metrics are exposed at /metrics (litellm/proxy/prometheus_metrics_server.py).
Architecture and concurrency. Pure Python + asyncio. The proxy is a FastAPI app deployed via uvicorn (ASGI server, configurable worker count via --num_workers, proxy_cli.py:295-329). Async at every layer: httpx.AsyncClient for outgoing provider calls, asyncio for concurrent request handling, and asyncpg via Prisma for DB access. Redis caching uses redis-py async. Deployment slots use Python anyio semaphores for max-parallel-requests enforcement (router.py:9035-9050).
Deployment modes. The CLI (litellm/proxy/proxy_cli.py) uses click to launch uvicorn with --port, --host, --num_workers, --max_requests_before_restart (for memory leak defense), --timeout_worker_healthcheck, and JSON logging support. Official Docker images (Dockerfile, docker-compose.yml) include PostgreSQL + Redis containers. Helm charts and Terraform scripts are provided for Kubernetes deployments. The --config flag points to a YAML/JSON config file defining models, routing, auth, and guardrail settings.
High availability. PostgreSQL is the durable backing store (via Prisma ORM with connection pooling via PgBouncer). Redis can be clustered for caching resilience. LiteLLM supports DB-native compaction, query-engine reaping, and graceful worker restarts. Logging is pluggable via litellm.callbacks — Langfuse, Datadog, New Relic, and custom loggers are first-class integrations.
litellm-rust/, loaded as litellm.rust_bridge._native) ships in the wheel; per-route rollout rules keep chat and embeddings on Python, let Anthropic messages opt in via LITELLM_RUST, and require Rust for OCR and Bedrock transcription.QuantumNous/new-api
answeredObservability. The project has no OpenTelemetry integration — there are no trace exporters, metric collectors, or OTel SDK references in the code. Observability is instead built on: (1) File-based logging (logger/logger.go) — SetupLogger() writes to a timestamped log file in the configured --log-dir, rotating at startup. The LogInfo/LogWarn/LogError/LogDebug functions write structured messages with context from the Gin request context. (2) Request logging (model/log.go) — every API request is recorded in the Log table (or ClickHouse log DB), storing tokens, quota, channel, latency (UseTime), IP, model, and extended billing metadata. This serves as the audit trail and usage analytics source. (3) Performance metrics (pkg/perf_metrics/) — relay latency and business rejection rates are captured at the controller level (controller/relay.go:134). (4) Active connection stats (middleware/stats.go) — an atomic counter tracks concurrent requests. (5) Admin dashboard — the React UI provides real-time usage dashboards, log exploration, and performance charts.
Performance architecture. The project is written in Go (compiled, with a green-thread concurrency model via gopool). It uses the Gin web framework, which is an async Go framework. The thread pool handles concurrent HTTP relay connections efficiently — each relay request is a lightweight goroutine. Redis offloads rate limiting state, channel polling indices, and cache lookups, reducing database pressure. The database layer supports ClickHouse as a separate log database (docker-compose.yml:32) to isolate analytics writes from transactional data.
Deployment modes. (1) Docker — the primary deployment path (Dockerfile) uses a multi-stage build: Bun builds the React frontend, Go compiles the binary, and a Debian slim image runs it on port 3000. The docker-compose.yml bundles the app with PostgreSQL (or MySQL) and Redis, with config via environment variables. (2) Standalone binary — go build produces a single static binary with embedded web frontend. (3) Electron desktop wrapper — electron/ provides a desktop app wrapping the same web UI. (4) systemd service (new-api.service) is included for Linux native deployment.
High availability. The architecture is single-node — there is no built-in clustering or HA mode. However, the system is stateless (session state in Redis, config in database), so multiple instances can share the same Postgres+Redis backend. The NODE_NAME env var (docker-compose.yml:38) identifies nodes in audit logs. The SESSION_SECRET must be shared across nodes (docker-compose.yml:41). Graceful connection handling is configurable via STREAMING_TIMEOUT and RELAY_IDLE_CONN_TIMEOUT.
Deployment configuration is entirely environment-variable driven: SQL_DSN, LOG_SQL_DSN, REDIS_CONN_STRING, TZ, BATCH_UPDATE_ENABLED, ERROR_LOG_ENABLED, plus session security settings (SESSION_SECRET, SESSION_COOKIE_SECURE, SESSION_COOKIE_TRUSTED_URL).
songquanpeng/one-api
answeredLanguage and concurrency. One API is written in Go using the Gin web framework (Go's analog of Express/Flask). Go's goroutine-based concurrency model and the Gin framework's non-blocking I/O provide good per-instance throughput. The server starts in main.go:103-120: a single Gin engine with middleware stack (session store, request ID, language, logger, rate limits) and route registration. No explicit worker count or multi-process configuration is visible — Gin runs as a single process with goroutine-based concurrency.
Logging. The middleware/logger.go sets up standard Gin request logging (method, path, status code, latency, client IP). Application-level logging uses common/logger/logger.go with log levels (Info, Error, Debug, SysLog). config.DebugEnabled and config.DebugSQLEnabled toggle verbose logging and SQL query logging. ONLY_ONE_LOG_FILE controls log file rotation (config.go:159).
Metrics and observability. No OpenTelemetry, Prometheus, or Datadog integration is present. The only built-in metric is the channel success-rate monitor (monitor/metric.go), which uses in-memory ring buffers per channel. There are no external metrics exporters.
Database. The system supports SQLite (default, file-based), MySQL, and PostgreSQL (model/main.go:67-109). Schema is auto-migrated via GORM's AutoMigrate on startup. A separate LOG_SQL_DSN can route usage logs to a secondary database. Connection pooling is configured via SQL_MAX_IDLE_CONNS, SQL_MAX_OPEN_CONNS, and SQL_MAX_LIFETIME environment variables (model/main.go:214-217).
Deployment. A multi-stage Dockerfile builds static Go binary + React frontends. The docker-compose.yml provides a full stack with Redis, MySQL, and the One API service. Environment variables control all configuration: SQL_DSN, REDIS_CONN_STRING, SESSION_SECRET, NODE_TYPE, SYNC_FREQUENCY, etc.
High-availability / multi-node. The NODE_TYPE environment variable (config.go:105) supports "master" vs "slave" nodes. Only the master node runs database migrations. Slave nodes can serve API traffic but delegate configuration reads to the database and rely on SYNC_FREQUENCY to refresh their in-memory channel cache. The FRONTEND_BASE_URL env var allows slave nodes to offload the React frontend to a separate server. There is no load balancer or clustering logic in the application itself — HA relies on an external LB + shared DB/Redis.
Health checks. The docker-compose defines a healthcheck at http://localhost:3000/api/status. The /api/status endpoint (controller/misc.go) provides a simple status probe.
Portkey-AI/gateway
answeredLogging is in-memory and SSE-streaming. The logHandler middleware (src/middlewares/log/index.ts:149) captures request/response data into a requestOptions array on the Hono context. After each request, processLog (line 112) sanitizes sensitive headers and broadcasts the log entry to all connected SSE clients at /log/stream. The log stream is protected by adminAuthMiddleware and includes timing, provider, status, and truncated response bodies (capped at 100KB, line 6).
Metrics and traces are minimal. The LogsService at src/handlers/services/logsService.ts:96-162 can construct OTLP-format span objects for tool execution events (with traceId, spanId, parentSpanId, and gen_ai.* attributes), but these are not exported to any OTel collector by the gateway — the structure is defined but the exporter is not implemented. The logger is plain console (src/apm/index.ts:1).
Performance architecture: Built on Hono (ultra-fast web framework designed for edge), written in TypeScript. Concurrency is handled by Hono's async middleware pipeline and Cloudflare Workers' waitUntil for fire-and-forget logging. There is no per-request thread pool — on Node.js it uses the async event loop, and on Workerd it uses the Workers runtime. Hono's sub-millisecond routing overhead is a key design goal (README claim, but the architecture supports it).
Deployment modes: Three runtimes are supported: Cloudflare Workers (npm run dev on wrangler), Node.js (npm run dev:node or npm run start:node), and any Hono-compatible runtime (Bun, Deno via adapter). The getRuntimeKey() check in src/index.ts and src/middlewares/log/index.ts switches behavior per runtime. The Node.js server (src/start-server.ts) uses @hono/node-server and supports WebSocket upgrades for /v1/realtime, SSE log streaming, and a local admin UI at /public/ that can be hidden with --headless. A conf.json file controls plugins, caching, and admin credentials.
High Availability: The open-source gateway itself has no built-in HA — deployments rely on the underlying platform (Cloudflare Workers' global network, Node.js process manager, Docker orchestrator). Redis provides a shared cache for multi-instance deployments.
higress-group/higress
answeredHigress provides extensive observability, deploys as a Docker container or via Helm on Kubernetes, and runs as a high-performance Envoy-based proxy.
Logs and traces: The ai-statistics plugin writes structured AI logs (ai_log) and OpenTelemetry-style trace spans. At the request header phase, it sets the span kind gen_ai.span.kind = LLM and extracts the route, cluster, API name — ai-statistics/main.go:700-750. On the response, it records model name, input/output/total tokens, llm_first_token_duration, and llm_service_duration — ai-statistics/main.go:1069-1098. Span attributes use the gen_ai.* prefix for ARMS (Alibaba Resource Monitoring System) compatibility — ai-statistics/main.go:102-108. Built-in attributes like question, answer, reasoning, tool_calls, system, reasoning_tokens, cached_tokens, and full input_token_details/output_token_details maps are extracted from both OpenAI and Anthropic response formats — ai-statistics/main.go:115-144. Streaming responses are re-assembled using an SSE framer for accurate token extraction — ai-statistics/main.go:895-919.
Metrics: The plugin defines Envoy-native counter metrics at the granularity of route.{route}.upstream.{cluster}.model.{model}.consumer.{consumer}.metric.{metricName} — ai-statistics/main.go:481-483. Metrics include token counts, request duration, stream duration count, and failure count. Error detection works across both streaming and non-streaming responses, checking response-level error fields and HTTP status codes — ai-statistics/main.go:1519-1560.
Performance architecture: Higress is built on Envoy (C++), with extension logic running as WASM plugins compiled from Go using TinyGo — plugins/wasm-go/Makefile:1-8. The WASM sandbox provides process-level isolation. Request body buffering is configurable per-extension (default 100MB) — ai-proxy/main.go:30. The WASM VM memory limit triggers automatic VM rebuild when exceeded — ai-endpoint-picker/main.go:230-238. Lua scripts in Redis are used for atomic rate-limit operations — ai-token-ratelimit/main.go:58-79.
Deployment modes: The quickstart uses Docker with included configuration (docker run -d --rm --name higress-ai) — README.md:70-72. Kubernetes deployment uses Helm charts and CustomResourceDefinitions for plugin configuration management — helm/. The API proto definitions at api/networking/v1/ define the control plane data models — api/networking/v1/http_2_rpc.proto.
HA and scaling: The failover mechanism (failover.go) uses VM-level leader election (CAS-based lease) to coordinate health checks across Wasm VMs — ai-proxy/provider/failover.go:292-341. Token failover tracks per-token health with configurable failure/success thresholds and cooldown recovery — ai-proxy/provider/failover.go:349-403. The x-higress-fallback-from header enables safe internal redirects across the gateway chain — ai-proxy/main.go:189-196.
maximhq/bifrost
answeredObservability: Bifrost has three observability layers. (1) Prometheus metrics via the telemetry plugin (plugins/telemetry/main.go) — exposes counters for upstream requests, success/error rates, input/output tokens, cache hits, and cost; histograms for upstream latency, overhead latency, stream first-token/inter-token latency, and retry counts. Supports push to a Prometheus Push Gateway for multi-node aggregation. (2) OpenTelemetry via the otel plugin (plugins/otel/main.go) — exports traces and metrics to one or more collector profiles (OTLP over HTTP or gRPC). Each profile has independent endpoints, headers, TLS config, and span filters. Traces follow the GenAI semantic convention and include per-provider spans, fallback spans, and overhead breakdowns. Supports distributed tracing via W3C traceparent and session-based trace grouping. (3) Structured logging via the logging plugin (plugins/logging/main.go) — logs all requests/responses to PostgreSQL with search, filter, and pagination via the logstore framework. Performance architecture: Written in Go with fasthttp (a high-performance async HTTP client/server), object pooling for responses, lock-free sync.Map data structures in hot paths, and per-provider connection pools with configurable concurrency. The model catalog uses memoization with generation counters to avoid recomputation. Deployment: Docker images (Alpine-based multi-stage builds embedding a React UI) run the HTTP server on port 8080 by default. Helm charts (helm-charts/bifrost/) deploy to Kubernetes with configurable replica count, secret management, and values examples for various backing stores (PostgreSQL+Weaviate, SQLite+Redis, etc.). The Dockerfile (transports/Dockerfile) builds a static Go binary using CGO. High availability: Horizontal scaling is supported by shared PostgreSQL and vector stores; the KV store (framework/kvstore/) provides distributed state for session affinity and rate-limit coordination across nodes. The OTel circuit breaker (plugins/otel/main.go:56-64) prevents a failed collector from blocking request processing.
katanemo/plano
answeredPlano is observed through OpenTelemetry tracing and structured logging, deployed as a Docker container with supervisord managing three processes, and scales horizontally behind a load balancer.
Logs, metrics, and traces: OpenTelemetry is the primary observability pillar. The tracing module in crates/brightstaff/src/tracing/ initializes a tracer (init.rs) with configurable sampling rate and OpenTelemetry gRPC endpoint (configuration.rs:480-485). Custom span attributes are injected from HTTP headers via collect_custom_trace_attributes(). LLM-specific semantic conventions are defined in tracing/constants.rs:49 (model name, provider, request/response content length, temperature, token counts). A dedicated PostHogExporter (tracing/posthog_exporter.rs:43) translates LLM spans into $ai_generation events and POSTs them to PostHog's batch/ API, supporting distinct_id from a configurable request header. The otlp_exporter gRPC target is also natively supported via Envoy's OpenTelemetry tracing integration (envoy.template.yaml:47-57).
Metrics are instrumented at two levels: WASM-level counters/gauges via proxy_wasm::hostcalls (crates/common/src/stats.rs) and application-level metrics collected throughout brightstaff (crates/brightstaff/src/metrics/). The time_to_first_token histogram is configured in envoy.template.yaml:7-28. Access logs are written to /var/log/access_ingress.log in standard Envoy format.
Performance architecture: The system is written in Rust (edition 2021, safe Rust throughout). The two WASM plugins (prompt_gateway, llm_gateway) are cdylib crates compiled for wasm32-wasip1 and run inside Envoy's WASM sandbox — meaning synchronous, callback-driven execution with no tokio/async, no networking, and no filesystem access. The native brightstaff binary runs as a separate async Tokio process for routing, state, and tracing. Communication between Envoy and brightstaff happens over HTTP; brightstaff exposes an orchestrator API (orchestrator_url) for routing decisions.
Deployment: A single Docker image bundles all components. supervisord.conf orchestrates three programs: config_generator (Python CLI, runs once to render YAML configs), brightstaff (native binary), and envoy (the proxy). The config generator runs envsubst on rendered templates to inject environment variables. The planoai Python CLI (cli/planoai/main.py) provides up, down, build, logs, trace and init commands for local development and deployment management.
High availability: Session bindings can be stored in Redis (crates/brightstaff/src/session_cache/redis.rs) rather than in-memory, enabling multiple brightstaff instances to share session state behind a load balancer. The Redis ConnectionManager with automatic reconnection handles connection drops. Session-cache lookups use short timeouts (connect: 5s, response: 2s) so a slow backend never blocks routing decisions. The project does not implement active health checks — it relies on Envoy's upstream cluster health monitoring and the response code fallback chain. Key rotation is handled through env-substituted configs and passthrough_auth mode.
tbphp/gpt-load
answeredLogging — the system uses logrus with structured fields. Request logs are written via telemetry.RequestLogSink (handler.go:485-510) — every request produces a requestRecorder that collects operation, model, cost, status, and timing. The requestlog package (internal/requestlog/) persists logs with configurable retention (default 7 days, maxRequestLogRetentionDays 365, internal/state/runtime_settings.go:59-61). Multiple log stores are supported: access-key usage, credential window usage, group usage, RPM history, and quota history.
Metrics — health.StatsStore (internal/health/stats.go:14) tracks per-credential success/failure/problem counts in sliding 5-minute windows (1-minute buckets). RPM counters are stored in rpm.Store and queried via ratelimit.AccessKeyRPM. No OpenTelemetry export is implemented — metrics are consumed in-process for health decisions and exposed via the control API.
Traces — no distributed tracing (OpenTelemetry or similar). The system uses request IDs (handler.newRequestID(), handler.go:464) that flow in X-GPTLoad-Request-ID response headers and appear in all log entries, enabling request correlation.
Deployment — packaged as a single Go binary via Docker (Dockerfile). The build uses multi-stage: Node.js for the web frontend, then Go compilation. The final image is Alpine-based. A docker-compose.yml and docker-compose.voice.yml are provided for container orchestration. The binary supports Windows as a first-class target with service management (service_windows.go, service_cli.go).
Concurrency model — Go language with goroutines (green threads). The gin HTTP framework provides non-blocking request handling. The concurrency.go package tracks active requests per access-key and globally. channel.RouteMode and the scheduler are pure in-memory operations with no per-request blocking IO beyond the upstream HTTP call.
Configuration management — state.Manager (internal/state/manager.go:11) holds an atomic pointer to the current ConfigSnapshot. Snapshots are immutable and replaced atomically on publish via sync/atomic.Pointer. The SnapshotReconciler interface synchronizes infrastructure resources before a snapshot becomes visible. Configuration is loaded from a database via state/loader and refreshed in-place without restarts.
High availability — the system is single-process; there is no built-in clustering, leader election, or shared state between instances. HA would rely on running multiple instances behind a load balancer with a shared database for configuration persistence. RPM and concurrency limits are in-process and not shared across nodes.
Health endpoints — the health package exposes credential health stats via the control API, and the auth layer checks credential cooldown/blacklist state inline during scheduling.
theagentrouter/agent-router
answeredThe project is observed via OpenTelemetry for both metrics and traces. Metrics are backed by OpenTelemetry with configurable exporters: Prometheus (always enabled), console exporter, and OTLP exporter. Environment variables OTEL_METRICS_EXPORTER, OTEL_EXPORTER_OTLP_ENDPOINT, and OTEL_SDK_DISABLED control configuration (internal/metrics/metrics.go:35-94). The Metrics interface tracks request start/completion timing, token usage, backend selection, time-to-first-token, and inter-token latency for streaming (internal/metrics/metrics.go:97-127). Traces support the same exporters (console, OTLP) with OTEL_TRACES_EXPORTER and auto-propagation via autoprop. Span recorders exist for every endpoint type (chat completions, embeddings, responses, speech, transcription, MCP, etc.) with semantic conventions for both OpenInference and OTEL GenAI (internal/tracing/tracing.go:137-328). Performance architecture: The Go-based ext_proc sidecar handles request/response transformations per-stream via gRPC. Router processors are tracked per-request-ID with a read-write mutex for concurrent access (internal/extproc/server.go:54-64). The data plane itself runs on Envoy Proxy (C++), providing battle-tested HTTP/2, connection pooling, and async I/O. Deployment modes: (1) Standalone CLI: aigw run downloads and configures an Envoy binary with the ext_proc sidecar as a local process (cmd/aigw/main.go:38-40). (2) Kubernetes: runs as a controller-runtime operator that reconciles Gateway resources, injects the ext_proc as a sidecar container (DaemonSet or init container), and manages rolling updates via workload template annotations (internal/controller/gateway.go:1298-1391). (3) Admin server: AdminPort exposes /metrics and /health HTTP endpoints (cmd/aigw/main.go:49). HA: Kubernetes controller pattern provides self-healing via reconciliation; Envoy's data plane handles connection retries and failover. Config bundle checksums ensure data-plane integrity (internal/filterapi/config_bundle.go:21-91).