LLMs Technical Reviews

elie222/inbox-zero

AI email assistant for Gmail and Outlook: plain-English rules triage new mail and pre-draft replies, plus a chat agent for the inbox.

GitHub ↗★ 12kTypeScriptAGPL-3.0commit b52de26 · 2026-10-06homepage ↗

Overview

Inbox Zero is a hosted-or-self-hosted AI assistant for one thing: email. It connects to Gmail or Outlook (plus Google or Microsoft calendars and Drive/OneDrive). It does its work in two ways.

The first is an automation pipeline that runs on every incoming message. You write rules in plain English (“newsletters: label and archive”, “anything needing a reply from me: draft one”). Gmail Pub/Sub or Microsoft Graph webhooks wake the app. A mix of static conditions, a cold-email detector and an LLM picks the matching rule, and the rule’s actions run: label, archive, mark read, draft or send a reply, forward, call a webhook, notify a Slack or Telegram channel, and more. Attachment filing runs alongside. Reply drafts pull in context from the thread, a knowledge base, past emails, the calendar, connected MCP servers and learned “reply memories” about how you edit drafts.

The second is a chat assistant, used on the web or from Slack and Telegram. It is a Vercel AI SDK ToolLoopAgent with about 25 tools for searching and reading mail, bulk inbox actions, editing rules, memories and settings. Anything that sends email, and some risky rule changes, come back as a pending card for you to confirm.

The codebase is a large pnpm/Turborepo monorepo (about 4,600 files). The core is the Next.js app in apps/web. A BullMQ worker, an Electron desktop shell and several shared packages sit around it. The root is AGPL-3.0, and apps/web/ee (billing) is under a separate commercial license. This is a mature product codebase, not a framework, so expect product logic everywhere.

Architecture

flowchart LR
  GM["Gmail Pub/Sub / Graph webhooks"] --> WH["Webhook routes"]
  WH --> PH["processHistoryItem"]
  PH --> RR["runRules"]
  RR --> MR["findMatchingRules"]
  MR --> LLM["LLM layer (AI SDK, fallback chains)"]
  RR --> EX["executeAct: label, archive, draft..."]
  EX --> EP["EmailProvider (Gmail / Outlook)"]
  UI["Web chat / Slack / Telegram"] --> CR["/api/chat route"]
  CR --> AG["aiProcessAssistantChat"]
  AG --> TLA["ToolLoopAgent + tools"]
  TLA --> LLM
  TLA --> EP
  DB[("Postgres via Prisma")] --- RR
  DB --- CR
Component Path Role
Web app apps/web/app Next.js App Router UI and API routes (webhooks, chat, cron, Slack/Telegram events, MCP server, mobile)
Webhooks apps/web/app/api/google/webhook, apps/web/app/api/outlook/webhook, apps/web/utils/webhook/ Acknowledge provider pushes, fetch history, dispatch each new message
Rules engine apps/web/utils/ai/choose-rule/ runRules, findMatchingRules, aiChooseRule, choose-args, executeAct
Drafting apps/web/utils/reply-tracker/, apps/web/utils/ai/reply/ Context gathering and draft generation, draft tracking, reply memories
Chat assistant apps/web/utils/ai/assistant/ System prompt, tool registry, memory tools, compaction, confirmation cards
LLM layer apps/web/utils/llms/ Provider list, model selection by tier, fallback, sensitive-data policy, prompt hardening
Email providers apps/web/utils/email/ EmailProvider interface with GmailProvider and OutlookProvider
Data apps/web/prisma/schema.prisma Accounts, rules, actions, executed rules, chats, memories, knowledge
Worker apps/worker/ BullMQ consumer that calls back into internal API routes
Desktop apps/desktop/ Electron shell
Packages packages/ Mail rendering (mail-core, mail-ui, mail-react), cli, scheduling, tinybird, transactional email

How a request flows

Take a new email arriving in a connected Gmail account.

  1. Push. Google Pub/Sub POSTs to /api/google/webhook. The route checks the verification token, decodes the history id, looks up the account and skips work if Gmail is rate-limiting. It acknowledges at once and processes in after() so Pub/Sub does not time out (apps/web/app/api/google/webhook/route.ts).
  2. History to messages. processHistoryForUser walks Gmail history since the last synced id. Each new message goes through the provider-neutral processHistoryItem.
  3. Pre-rule work. The shared item processor optionally categorises the sender (used by category filters in rules). It then calls runRules when the account has automation rules and AI access (apps/web/utils/webhook/process-history-item.ts). Attachment filing runs in parallel in after().
  4. Match. runRules splits ordinary rules from conversation-tracking meta rules and calls findMatchingRules (apps/web/utils/ai/choose-rule/run-rules.ts). That function runs the cold-email check first. It evaluates static conditions (from/to/subject globs, groups, categories), and sends only rules with AI conditions to an optional decision model or to aiChooseRule, with past classification feedback for that sender (apps/web/utils/ai/choose-rule/match-rules.ts).
  5. Arguments. For each matched rule, executeMatchedRule drops actions that low-trust static from patterns may not trigger. It then fills templated action fields with AI-generated arguments, for example the body of a draft (apps/web/utils/ai/choose-rule/run-rules.ts).
  6. Execute. executeAct runs each action item through the EmailProvider. It skips automated archive for protected senders and records a per-action outcome (apps/web/utils/ai/choose-rule/execute.ts). Draft actions use the drafting pipeline, which gathers knowledge, reply memories, history, calendar availability, writing style and MCP context in parallel (apps/web/utils/reply-tracker/generate-draft.ts).
  7. Record. An ExecutedRule row stores the match reason and actions. It drives the history view and the “fix this rule” flow in chat.

The chat path is shorter. POST /api/chat (apps/web/app/api/chat/route.ts) loads the chat, compacts old history if needed, and loads the 20 newest ChatMemory rows (apps/web/app/api/chat/route.ts). aiProcessAssistantChat then streams a ToolLoopAgent run back as a UI message stream.

Key components

Rules as data

A rule is a Prisma row with static conditions, an optional natural-language AI condition and a list of typed actions. Each action has templated fields that the LLM can fill. Deterministic matching always runs first, and only rules with AI conditions reach the model. That keeps cost down and makes most matches explainable. “Conversation” meta-rules (To Reply, FYI, Awaiting Reply) track thread state across messages, which is how the Reply Zero feature works.

The LLM layer

utils/llms/ wraps the Vercel AI SDK. Providers include OpenAI, Azure, Azure Foundry, Vertex, Anthropic, Bedrock, Google, Groq, Cerebras, OpenRouter, Vercel AI Gateway, Ollama, any OpenAI-compatible endpoint, and the Codex CLI and Claude Code subscriptions (apps/web/utils/llms/config.ts). Models are picked by tier from DEFAULT_LLMS, ECONOMY_LLMS, CHAT_LLMS, NANO_LLMS and DRAFT_LLMS. Each is an ordered provider:model chain where later entries are fallbacks (apps/web/env.ts). Users can also bring their own key. Every call goes through prompt hardening and a per-account sensitive-data policy (ALLOW/REDACT/BLOCK) that covers both messages and tool outputs (apps/web/utils/llms/index.ts).

Chat assistant and confirmation

The tool registry is assembled per request. Send, reply and forward tools exist only when email sending is enabled, and calendar tools only when a calendar is connected (apps/web/utils/ai/assistant/chat.ts). Those tools do not send anything. They return a requiresConfirmation: true, confirmationState: "pending" payload (apps/web/utils/ai/assistant/chat-inbox-tools.ts). The UI renders it as a card, and only confirmAssistantEmailActionForAccount performs the send after the user clicks (apps/web/utils/actions/assistant-chat-confirmation.ts). The agent also has a time budget: 720 s on the web and 60 s from Slack or Telegram (apps/web/utils/ai/assistant/chat.ts). After that, prepareStep switches off all tools and forces a text answer (apps/web/utils/ai/assistant/chat.ts).

Memory

There are three separate stores. ChatMemory holds facts the chat agent saves. The newest 20 are injected as a context message each turn (apps/web/utils/ai/assistant/chat.ts). Facts sourced from the user must quote verbatim evidence, and inferred ones need confirmation. ReplyMemory (facts, preferences, procedures) is extracted from the difference between an AI draft and what you actually sent, then selected per email for future drafts. A summarised learned writing style is derived from it too. Knowledge is a user-curated knowledge base used in drafting. All three are plain Postgres rows with no vector store.

Extending it

  • Rules and prompts. Most behaviour is configured, not coded: rules, personal instructions, the knowledge base, the cold-email prompt and digest settings, all editable in the UI or through the chat agent.
  • MCP. The app can connect to external MCP servers whose tools feed draft context (utils/ai/mcp/). It also exposes its own MCP server (app/api/mcp-server) so other AI clients can work with Inbox Zero data.
  • API and messaging. app/api/v1 is an external API with API keys. Slack and Telegram event routes carry the same assistant into those apps.
  • New email provider. Implement the EmailProvider interface used by createEmailProvider (apps/web/utils/email/provider.ts). This is a large surface, because Gmail and Outlook semantics (labels vs folders, history vs delta) leak into many features.

Running it

  • Services. Postgres 16 and Redis are required. The root docker-compose.yml defines db, redis, serverless-redis-http (an Upstash-compatible HTTP front for Redis), web, an optional worker profile and a cron container. The cron container curls internal endpoints with CRON_SECRET: scheduled actions, automation jobs, follow-ups, digests, meeting briefs, meeting-recorder scheduling and watch renewal (docker-compose.yml).
  • Accounts. You need Google OAuth credentials and a Pub/Sub topic for Gmail push, or a Microsoft app registration for Outlook, plus at least one LLM provider key (or a local Ollama/OpenAI-compatible endpoint). Billing, analytics and error tracking are optional.
  • Image. docker/Dockerfile.prod builds the web app on node:24-alpine with pnpm. A setup CLI lives in packages/cli.

Strengths and caveats

  • Strength: deterministic before LLM. Static conditions, sender categories and cold-email guards filter mail before any model call. Every decision is stored as an ExecutedRule with a reason, so automation can be audited and corrected.
  • Strength: sending stays behind a human. Chat-initiated sends, replies and forwards are confirmation cards, and send tools can be disabled entirely by env.
  • Strength: drafts that learn. Reply memories mined from your edits, a learned writing style and multi-source context make drafting more than one prompt.
  • Strength: broad model support. Many providers, tiered model chains with fallbacks, local models and a sensitive-data redaction layer.
  • Caveat: rule actions are not confirmed per message. Once a rule is active, its actions run unattended, including REPLY, SEND_EMAIL, FORWARD, DELETE and CALL_WEBHOOK. The only risk check is when a rule is created or changed through chat.
  • Caveat: heavy to self-host. Postgres, Redis, OAuth app setup, Pub/Sub or Graph subscriptions with periodic renewal, and a cron sidecar. Many optional SaaS hooks are woven through the code.
  • Caveat: email-only scope. Calendar, Drive and meeting features all serve the inbox. It is not a general assistant.
  • Caveat: size. Core files like match-rules.ts (1,104 lines) and llms/index.ts (2,121 lines) carry a lot of product branching. Contributing needs real familiarity with the codebase.

Sources: code at b52de26, verified Q&A.

How it answers the Open-source personal assistants questions

Each answer was drafted by a code-reading agent at commit b52de26. Its citations were checked mechanically. Compare with the other open-source personal assistants →

How is the assistant architected?

answered

Agent loop and runtime. The user request flows through the Next.js API route at apps/web/app/api/chat/route.ts, which receives messages via POST. The route validates the input, loads the chat + compaction history from PostgreSQL, optionally compacts old messages, loads up to 20 ChatMemory entries, then calls aiProcessAssistantChat (apps/web/utils/ai/assistant/chat.ts). This function builds a system prompt, constructs a ToolLoopAgent from the Vercel AI SDK (via toolCallAgentStream in apps/web/utils/llms/index.ts:886), and streams tool calls and text back as a UIMessageStream. The agent runs in a loop — each step can call one tool, the result is fed back, and the loop continues until the model decides to respond or the tool budget expires (720s for web, 60s for messaging).

Frontend/backend split. The frontend is a Next.js app-router server-rendered React application with Tailwind CSS and shadcn/ui components. Data is fetched client-side using SWR hooks; mutations go through next-safe-action server actions (e.g., confirmAssistantEmailAction in apps/web/utils/actions/assistant-chat.ts:15). POST API routes are reserved for mobile-native integrations.

Main packages. The monorepo (pnpm-workspace.yaml) contains: apps/web (Next.js), apps/desktop (Electron shell), apps/worker (background queue consumer), packages/mail-core/mail-react/mail-ui (email rendering), packages/api (shared API types), packages/loops (automation jobs), packages/scheduling (booking/availability), packages/tinybird (analytics).

Request-to-action flow. The user types in the chat UI → POST to /api/chat → messages converted to ModelMessage[] → ToolLoopAgent runs with ~25 tools (searchInbox, manageInbox, sendEmail, etc.) → each tool calls the Google Gmail or Microsoft Graph API via the GmailProvider/OutlookProvider classes — tools that change state (send, delete rule, save inferred memory) return requiresConfirmation: true and confirmationState: 'pending' (e.g., apps/web/utils/ai/assistant/tools/rules/create-rule-tool.ts:69-75), which the frontend renders as a confirm button that fires a server action.

How are integrations (email, calendar, chat, docs) implemented?

answered

Supported services. The app integrates with Google/Gmail (Gmail API, Google People API, Google Calendar) and Microsoft Outlook (Microsoft Graph API for both mail and calendar). It also connects to Slack and Discord as messaging channels via MessagingChannel (apps/web/prisma/schema.prisma:1700), and Google Drive/OneDrive via DriveConnection (apps/web/prisma/schema.prisma:1800). There is no MCP integration for external services in the chat assistant — it's all native API clients.

API clients vs MCP. The app uses the official Google and Microsoft SDKs directly. GmailProvider (apps/web/utils/email/google.ts:131) wraps @googleapis/gmail with OAuth token refresh. OutlookProvider (apps/web/utils/email/microsoft.ts:141) wraps the Microsoft Graph API. Calendar integrations use GoogleCalendarEventProvider and MicrosoftCalendarEventProvider (apps/web/utils/calendar/providers/). There is also MCP server support (an MCP server that exposes Inbox Zero data to external AI clients) and MCP connections the assistant can configure, but the assistant's own tools do not use MCP.

OAuth flow. OAuth tokens are obtained through NextAuth.js-based flows. For Google, the getGmailClientWithRefresh function (apps/web/utils/gmail/client.ts:52) creates an OAuth2 client, sets credentials from stored tokens, and automatically refreshes the access token when it expires (10-minute buffer). Tokens are saved to the Account model (next-auth tables) via saveTokens (apps/web/utils/auth/save-tokens).

Token storage. OAuth access and refresh tokens are stored in the Account Prisma model under fields access_token, refresh_token, expires_at (apps/web/prisma/schema.prisma:22-26). Calendar connections store their own tokens in CalendarConnection.accessToken/refreshToken/expiresAt (apps/web/prisma/schema.prisma:1416-1434). Drive connections similarly in DriveConnection. All token storage uses the database.

Sync vs on-demand. The email client fetches on-demand via the Gmail/Outlook API — searchInbox, readEmail tools call the live API. The app also supports webhook-based push notifications (Google PubSub, Microsoft webhooks) and the watchEmailsExpirationDate field tracks subscription expiry. Background sync runs via worker jobs and cron.

How is memory and user context stored and retrieved?

answered

Database storage. Memories are stored in the ChatMemory PostgreSQL table (apps/web/prisma/schema.prisma:1401-1414), keyed by emailAccountId with an optional chatId link. The schema stores: content (the memory text), createdAt, updatedAt. There is no vector store — memory search uses PostgreSQL CONTAINS with case-insensitive matching on the content field (apps/web/utils/ai/assistant/chat-memory-tools.ts:35-48). A query of empty string returns the 10 most recent memories.

What is remembered. The LLM is prompted to save explicit user preferences and facts stated in chat. Two sources exist: user_message (user directly states a fact) and assistant_inference (model infers from context). The saveMemoryTool (apps/web/utils/ai/assistant/chat-memory-tools.ts:125-219) validates that user_message sources have matching evidence quotes from conversation history, enforced by validateUserMemoryEvidence (apps/web/utils/ai/assistant/chat-memory-policy.ts:5-78). Inferred memories (assistant_inference) require UI confirmation before being persisted. Memories are not applied to email processing rules — only assistant chat conversations.

Prompt injection. At the start of each chat turn, the route loads up to 20 recent memories (apps/web/app/api/chat/route.ts:265-278) and passes them into the system context as "Memories from previous conversations: [date] content" (apps/web/utils/ai/assistant/chat.ts:254-261). The LLM can also call searchMemoriesTool dynamically to recall specific facts.

Summarization/compaction. When a conversation exceeds ~8K tokens, compactMessages (apps/web/utils/ai/assistant/compact.ts:63) uses the LLM to summarize the history into a single text message. The compaction summary is stored in ChatCompaction (apps/web/prisma/schema.prisma:1387-1399) and injected as a system message in future turns. Compaction also triggers extractMemories (apps/web/utils/ai/assistant/compact.ts:152) which calls the LLM to identify and save durable facts from the conversation history as ChatMemory records.

Editor's note. Correction: memory also shapes email processing. ReplyMemory records (facts, preferences, procedures) are learned from how the user edits AI drafts before sending and are selected into future reply drafts, together with a learned writing style (apps/web/utils/ai/reply/reply-memory.ts).

How are actions on the user's behalf gated?

answered

Confirmation/human-in-the-loop. The primary safety mechanism is a confirmation gate on all state-changing operations. The manageInboxTool (apps/web/utils/ai/assistant/chat-inbox-tools.ts) supports archive_threads, trash_threads, label_threads, bulk_archive_senders, unsubscribe_senders, etc. Destructive tools like sendEmail, replyEmail, and forwardEmail (gated by NEXT_PUBLIC_EMAIL_SEND_ENABLED) are registered conditionally in the toolset (apps/web/utils/ai/assistant/chat.ts:300-305). Before execution, the tool returns { requiresConfirmation: true, confirmationState: 'pending' } — this is rendered as a "card" in the UI. The user must click a confirm button which calls a server action (confirmAssistantEmailAction in apps/web/utils/actions/assistant-chat.ts:15). The confirmAssistantEmailActionForAccount function (apps/web/utils/actions/assistant-chat-confirmation.ts:69) then executes the real API call.

Rule creation gate. createRuleTool (apps/web/utils/ai/assistant/tools/rules/create-rule-tool.ts:65-76) checks actionsNeedChatRiskConfirmation before every rule creation. If the rule contains potentially risky actions (like auto-archiving or auto-sending), it returns requiresConfirmation: true with warning messages. Low-risk rules execute immediately. Delete and update rule tools follow the same pattern.

Memory write gate. The saveMemoryTool only auto-saves source: "user_message" memories after validating the evidence quote exists verbatim in the conversation. Inferred memories (source: "assistant_inference") always return requiresConfirmation: true (apps/web/utils/ai/assistant/chat-memory-tools.ts:152-164).

Sensitive data policy. The app has a sensitiveDataPolicy field on EmailAccount with values ALLOW/REDACT/BLOCK (apps/web/prisma/schema.prisma:179). The LLM layer enforces this: enforceSensitiveDataPolicy wraps messages and wrapToolsWithSensitiveDataPolicy wraps tool outputs, stripping or blocking PII before they reach the model (apps/web/utils/llms/index.ts:298-311). A default policy is configured server-side via SENSITIVE_DATA_POLICY_DEFAULT env var.

Tool budget / timebox. The assistant enforces a running tool execution budget — 720s for web, 60s for messaging (apps/web/utils/ai/assistant/chat.ts:57-60). After the budget expires, prepareStep in the ToolLoopAgent disables all tools and forces a direct text response (apps/web/utils/ai/assistant/chat.ts:370-381).

Editor's note. Correction: not every state-changing chat tool is confirmation-gated. Send, reply and forward, inferred memories and risky rule changes return a pending card, but manageInbox (archive, trash, label, bulk-archive senders, unsubscribe) runs immediately. Automation-rule actions such as REPLY, SEND_EMAIL, FORWARD and DELETE also run unattended once a rule is active.

How are LLM providers selected and configured?

answered

Supported providers. The app supports 14 LLM providers configured via environment variables in apps/web/env.ts:7-23: anthropic, azure, azure-foundry, vertex, google, openai, bedrock, openrouter, groq, cerebras, aigateway, ollama, openai-compatible, codex-cli, claude-code. Each maps to a Vercel AI SDK provider (@ai-sdk/anthropic, @ai-sdk/openai, @ai-sdk/google, etc.) instantiated in apps/web/utils/llms/model.ts. The provider enum is in apps/web/utils/llms/config.ts:3-19.

Configuration surface. The primary configuration is via env vars: DEFAULT_LLMS, ECONOMY_LLMS, CHAT_LLMS, NANO_LLMS, and DRAFT_LLMS (apps/web/env.ts:116-120). Each is a colon-separated chain of provider:model entries — the first valid entry is the primary, later entries serve as fallbacks. Individual API keys are set per provider (ANTHROPIC_API_KEY, OPENAI_API_KEY, GOOGLE_API_KEY, etc.). Users can also set a personal aiApiKey, aiProvider, and aiModel on their User record (apps/web/prisma/schema.prisma:94-96), which overrides the deployment defaults. The per-use-case model type mapping (apps/web/utils/llms/use-cases.ts) routes LlmUseCase.AssistantChat to the "chat" model tier.

Tool-calling / structured outputs. All LLM interactions use the Vercel AI SDK (ai package). The main chat agent uses ToolLoopAgent (apps/web/utils/llms/index.ts:959) with a tool registry of ~25 Zod-defined tools. The generateObject wrapper adds JSON repair via jsonrepair on parse failures (apps/web/utils/llms/index.ts:168-178). Structured output is requested through the SDK's generateObject which uses provider-native structured output modes when available (OpenAI response_format, Gemini response_mime_type, etc.).

Local model support. Ollama is supported: the OLLAMA_BASE_URL env var configures the endpoint and OLLAMA_MODEL picks the model. The generic openai-compatible provider extends this to any OpenAI-compatible API endpoint, letting users run local models through tools like LocalAI or LM Studio (apps/web/utils/llms/model.ts:296-312).

Fallback chain. The selectDeploymentModelByType function builds a chain of model candidates from the env var. If the primary provider fails (rate limit, server error, content filter), the system falls through to the next provider in the chain (apps/web/utils/llms/index.ts:927), attempting each until one succeeds.

How is it deployed and self-hosted?

answered

Runtime dependencies. The app requires PostgreSQL 16 (main database), Redis 7 (job queues via BullMQ or Upstash REST API for rate-limiting/caching/subscriptions), and optionally an SMTP/email sending provider. The docker-compose.yml (apps/web/../docker-compose.yml) defines db (Postgres), redis, serverless-redis-http (HTTP Redis proxy for Upstash-compatible clients), web (Next.js application), worker (background queue consumer), and cron (scheduled jobs using Alpine curl hitting the app's internal cron API endpoints).

Docker/one-click paths. A production Docker image is built from docker/Dockerfile.prod using node:24-alpine with pnpm 11.19, running the monorepo without Electron. The web service runs next start (standalone mode). The worker runs start-worker.sh. The cron container fires 7 cron-like loops covering scheduled actions, automation jobs, follow-up reminders, digest emails, meeting briefs, meeting recorder scheduling, and email watch renewal — all via curl to internal API endpoints authenticated by CRON_SECRET. The docker-compose.local.yml variant supports local development with the full stack.

Required external accounts. The app needs Google OAuth credentials (GOOGLE_CLIENT_ID, GOOGLE_CLIENT_SECRET) and/or Microsoft credentials for email/calendar. For LLM functionality, at least one provider API key is required (e.g., ANTHROPIC_API_KEY, OPENAI_API_KEY). For production use: Google PubSub topic for push email notifications, Stripe/Lemon Squeezy for billing, Sentry for error tracking, PostHog for analytics, OpenRouter for router-style LLM access, QStash for serverless queue, and Vercel KV for distributed caching. The env.ts file documents every env var with Zod validation (apps/web/env.ts).

Self-hosted deployment. The recommended self-hosted path is Docker Compose. A setup CLI is available via pnpm setup (packages/cli/src/main.ts). The CRON_SECRET must be set for background job auth. A local setup .env.example documents all required and optional variables. The desktop app is an Electron shell (apps/desktop/package.json) for macOS/Windows/Linux that wraps the web app; it ships via electron-builder.

Vercel deployment. The app is deployable on Vercel with specific env vars for preview deployments (auto-base-url detection from VERCEL_URL), and uses Vercel KV as an alternative Redis backend.