LLMs Technical Reviews

How is a tool call executed?

Proxy vs direct; sandboxing; rate limiting; retries; pagination; error mapping.

Verdict

ACI runs every call through a central proxy. execute_function checks the app configuration, the agent’s allowed apps, the enabled functions and the linked account. It resolves and refreshes credentials and can ask an LLM to check the call against custom instructions. Then the matching executor runs. A REST executor validates the input against the visible schema, adds the hidden defaults and auth, and sends one httpx request (10 s connect, 30 s read). Errors become success: false results. There is a per-IP in-memory rate limit and per-project quotas, but no retries, no pagination helpers and no sandbox. The HTTP call also blocks inside an async route.

In Klavis, the Strata router is a pass-through. execute_action merges the JSON-string path, query and body params and calls session.call_tool on the downstream MCP server. If the server reports an error, the router returns it as text. The router has no retries, rate limits or call timeout; isolation comes only from running each server as its own process or container. Each server handles pagination and error mapping in its own way.

Choose ACI if you need central policy, auditing and uniform error results. Choose Klavis if you want a thin router and accept that behaviour differs from server to server.

More projects in this category are being researched.

Per-project answers

Klavis-AI/klavis

answered

Layered proxy. Strata (tools.py:246-286): execute_action looks up MCPClient, merges params, calls client.call_tool(). MCPClient (mcp_proxy/client.py:97-122): session.call_tool() on MCP ClientSession; result.isError raises RuntimeError. Transport (mcp_client_manager.py:82-118): HTTPTransport or StdioTransport. Base (transport/base.py:32-56) uses AsyncExitStack. Direct client (mcp_client.py:331-386): find_server_for_tool(), 30s progress updates. Error mapping: exceptions caught/logged, returned as TextContent (tools.py:288-292). No sandboxing (process/container isolation). No rate limiting/retries. Sensitive params redacted (mcp_client.py:109-139). Pagination is server-specific.

aipotheosis-labs/aci

answered

Function execution is a multi-step pipeline triggered at POST /v1/functions/{function_name}/execute. The execute_function in functions.py (318-486) runs these stages: (1) Lookup — fetches the Function DB record by name. (2) Authorization chain — validates the app's AppConfiguration exists and is enabled, checks the function's app is in the agent's allowed_apps, verifies the function is in the enabled list, and confirms the LinkedAccount exists and is enabled (functions.py:364-433). (3) Credential resolution — calls security_credentials_manager.get_security_credentials which handles OAuth2 token expiry/refresh transparently (security_credentials_manager.py:33-49, 87-139). (4) Custom instruction validation — runs the function input against the agent's custom_instructions via GPT-4o (functions.py:451-456). (5) Execution — selects an executor via get_executor(function.protocol, linked_account) (function_executors/init.py:21-32). For REST protocol, RestFunctionExecutor._execute constructs an httpx request by combining the server_url, path, path params, query, header, cookie, and body from the function_input, then injects credentials and sends it (rest_function_executor.py:37-103). No sandboxing is applied — REST calls go directly to third-party APIs with a 10s/30s timeout. For the CONNECTOR protocol, ConnectorFunctionExecutor dynamically imports and instantiates a Python class from server/app_connectors/ (e.g., Gmail, E2B, Vercel) via importlib (connector_function_executor.py:31-61). BaseExecutor validates input against the JSON Schema and injects invisible default values before delegation (base_executor.py:32-82). Rate limiting is IP-based via RateLimitMiddleware (ratelimit.py:19-72) with per-second and per-day limits using a moving window. Daily/monthly quotas are enforced per project in validate_project_quota and validate_monthly_api_quota (dependencies.py:86-158). Results and errors are logged with structured telemetry to Logfire, including input/output truncation at 8KB (functions.py:239-273). There is no explicit retry or pagination mechanism in the executor — pagination is delegated to each API's own parameters.

Editor's note. Correction: the custom-instruction check defaults to gpt-4o-mini, not GPT-4o (custom_instructions.check_for_violation).

← How are integrations exposed to LLM agents? · How are data sync, webhooks and triggers implemented? →