khoj-ai/khoj
Self-hosted multi-user chat server that answers from your pgvector-indexed notes, web search, sandboxed code and a computer-use agent.
Overview
Khoj is a self-hostable “second brain”: a multi-user chat server that answers from your own documents, the web, a code sandbox and, optionally, a remote-controlled desktop. You upload or sync notes (Markdown, Org, PDF, DOCX, plaintext, images, Notion pages, GitHub repos). Khoj chunks and embeds them into PostgreSQL with pgvector. A chat turn then picks which sources to consult, gathers context, and streams a cited answer from OpenAI, Anthropic or Gemini models, or any OpenAI-compatible local endpoint.
It is a server product more than a library. One Python process runs FastAPI for the API and chat, mounts a Django app for the ORM, migrations and admin panel, and serves a prebuilt Next.js web client. The admin panel is the main configuration surface: chat models, API keys, search model, web scraper and server-wide settings are database rows, not config files. Around the chat loop sit the “assistant” features: custom agents with their own persona, model and tool set; automations that re-run a query on a cron schedule and email the result; a long-term memory of facts extracted from conversations; and a /research mode that iterates over tools for several rounds.
It suits people who want a private Perplexity-plus-notes assistant on their own box, or a small team that wants shared agents over shared documents. Besides the web client, the repo ships Obsidian, Emacs, desktop and Android clients that talk to the same HTTP API. It does not run long autonomous tasks on your machine: the only “computer use” happens inside a sandboxed container.
Architecture
flowchart LR
C["Web client / API caller"] --> API["FastAPI routers (/api/*)"]
API --> AUTH["Auth middleware + rate limiters"]
AUTH --> EG["event_generator (chat turn)"]
EG --> SEL["Tool selection (LLM or /command)"]
SEL --> DOCS["Document search (pgvector)"]
SEL --> WEB["Web search + page reader"]
SEL --> CODE["Code sandbox (Terrarium / E2B)"]
SEL --> OP["Operator (computer / browser)"]
SEL --> RES["Research loop + MCP tools"]
EG --> GEN["agenerate_chat_response"]
GEN --> LLM["OpenAI / Anthropic / Gemini"]
EG --> SAVE["save_to_conversation_log"]
SAVE --> MEM["Memory extraction"]
SAVE --> PG["PostgreSQL + pgvector"]
DOCS --> PG
SCHED["APScheduler (automations)"] --> EG
| Component | Path | Role |
|---|---|---|
| Entry point | src/khoj/main.py |
Django setup and migrations, FastAPI app, scheduler leader election, route and static mounts |
| Route wiring | src/khoj/configure.py |
Registers routers, auth middleware, periodic jobs (schedule library) |
| Chat API | src/khoj/routers/api_chat.py |
HTTP and WebSocket chat endpoints, event_generator turn orchestration |
| Helpers | src/khoj/routers/helpers.py |
Command parsing, model fallback wrapper, response generation, memory updates, automations, content indexing |
| Research mode | src/khoj/routers/research.py |
Multi-iteration tool loop with parallel tool execution and MCP tools |
| Tools | src/khoj/processor/tools/ |
Online search and webpage reading, sandboxed Python, MCP client |
| Operator | src/khoj/processor/operator/ |
Vision-model agent that drives a Docker desktop or a browser |
| Conversation backends | src/khoj/processor/conversation/{openai,anthropic,google}/ |
Provider-specific chat and structured-output calls |
| Content ingestion | src/khoj/processor/content/ |
Per-format parsers that turn files into Entry rows |
| Data model | src/khoj/database/models/__init__.py, database/adapters/ |
Django models and query adapters (users, chat models, entries, memories, agents) |
| Web client | src/interface/web/ |
Next.js app built with Bun and served as static files |
| Other clients | src/interface/{obsidian,emacs,desktop,android}/ |
Editor plugins and apps that sync files and chat over the HTTP API |
How a request flows
Take a plain chat message, “what did I write about the Q3 budget?”, sent to POST /api/chat:
- Admit. The endpoint is wrapped in
@requires(["authenticated"])and three FastAPI dependencies: 20 requests per minute, a daily cap (100 free, 600 subscribed), and an image size limiter. It then hands the body toevent_generatorand either streams the output or collects it into one JSON response (api_chat.py). The WebSocket variant at/api/chat/wsfirst rejects origins outside the allowed hosts (api_chat.py). - Parse a command.
get_conversation_commandmaps a leading/notes,/online,/webpage,/code,/research,/image,/diagramor/operatorto a command. Anything else isDefault(helpers.py). - Load memory. If memory is enabled, the turn pulls recent memories (last 7 days) and semantically similar ones from
UserMemory(adapters). - Choose tools. For
Default, one LLM call (aget_data_sources_and_output_format) returns the sources and the output mode. On failure it falls back to a general, tool-less answer (api_chat.py). Per-command rate limits are then checked. - Gather context.
event_generatorruns the chosen tools one after another: research, document search, online search, webpage reading, code, operator, then image or diagram output. Each tool yields status events to the client. - Answer.
agenerate_chat_responsebuilds the final prompt from references, online and code results, memories and history, then dispatches onchat_model.model_typetoconverse_openai,converse_anthropicorconverse_gemini. - Persist and learn.
save_to_conversation_logappends both messages, with their context, to the conversation’s JSON log. Unless the turn came from an automation, it then callsai_update_memories(utils.py). That function asks the LLM which facts to create or delete, and embeds new ones intoUserMemory(helpers.py).
Key components
Model routing and fallback
A ChatModel row has a name, a type (openai, anthropic or google), a price tier, a vision flag and an optional AiModelApi with key and base URL (models). The base URL is what makes Ollama, vLLM or LM Studio work: they are configured as “openai” models. ServerChatSettings has six slots (default, advanced, and fast/deep variants for free and paid users), a web scraper choice, a priority and a server-level memory mode (models). send_message_to_model_wrapper tries the primary model, swaps to a vision model if images are attached, then walks the same slot across lower-priority settings rows. It moves on only for retryable errors (helpers.py).
Document search
Each upload becomes Entry rows with a pgvector embedding. The embedding model is a local sentence-transformers model by default, or a remote inference endpoint. Search applies file, date and word filters, ranks by cosine distance with a configurable threshold, and can re-rank with a cross-encoder. Notion (OAuth) and GitHub (personal access token) are content sources fed into the same pipeline. A 22-25 hour job re-indexes them under a process lock (configure.py).
Web search
search_online builds a provider list from whatever keys are set: Serper, Exa, Firecrawl, Google Custom Search, then SearXNG. It uses the first provider that returns results (online_search.py). Page reading uses the configured WebScraper type: Firecrawl, Olostep, Exa or direct fetch (models).
Research mode
research() runs up to KHOJ_RESEARCH_ITERATIONS (default 5) rounds. In each round apick_next_tool lets the model choose tool calls. The operator runs on its own because it streams. Everything else runs concurrently with asyncio.gather. A user can interrupt mid-run with a new instruction, which is folded into the history (research.py). MCP servers registered in the admin panel are exposed only here, as server/tool names. MCPClient uses SSE for http(s) paths and stdio for script paths (mcp.py).
Code and operator sandboxes
Generated Python runs in a Terrarium container (no network) or, with an API key, in an E2B sandbox. The operator is off unless KHOJ_OPERATOR_ENABLED is set (helpers). It needs a vision model. It uses the Anthropic computer-use agent, or a UI-TARS grounding agent; the OpenAI operator branch is hard-disabled. It drives a Docker desktop or a Playwright browser for up to 100 steps. A RequestUserAction ends the loop and hands the question back to the user (operator).
Automations
An automation is an APScheduler cron job stored in Postgres via DjangoJobStore. Its id is derived from the user and query hash. It runs scheduled_chat under a process lock, with 60 seconds of jitter (helpers.py). Only one worker executes jobs: main.run takes a ProcessLock to become schedule leader and starts the scheduler paused on the others (main.py).
Extending it
- Add a model. Create an
AiModelApi(key plus base URL) and aChatModelin the admin panel, then put it in aServerChatSettingsslot. No code change is needed for any OpenAI-compatible server. - Add tools via MCP. Register a
McpServer(name, path or URL, optional key). Its tools appear in research mode. - Custom agents. An
Agentbundles a persona, a chat model, input tools, output modes and a privacy level, and can own its own knowledge files. Agents are created through/api/agentsor the web UI. - New content types. Add a processor under
processor/content/that producesEntryobjects throughTextToEntries, then wire it intoconfigure_content. - New routes.
configure_routesis a plain list ofinclude_routercalls. Billing and Twilio routers are added conditionally (configure.py).
Running it
- Docker Compose. The stock compose file runs
pgvector/pgvector:pg15, the Terrarium sandbox, SearXNG, an optionalkhoj-computerdesktop for the operator, and the server image. The server needs only Postgres credentials, a Django secret and an admin email and password. Add at least one LLM key, such asOPENAI_API_KEY,ANTHROPIC_API_KEYorGEMINI_API_KEY, or point a model at a local endpoint. - pip.
pip install khojgives akhojcommand (khoj.main:run) with--host,--port,--socket,--anonymous-modefor single-user use without login, and--non-interactive(cli.py). You still need your own Postgres with pgvector. - Production.
prod.Dockerfileruns gunicorn and adds the Stripe, Twilio and S3 extras. Email login and automation emails use Resend.
Strengths and caveats
- Strength: complete product. Retrieval, web search, code, memory, agents, automations, admin panel and a polished web client ship together and work multi-user.
- Strength: model portability. Three native SDKs, plus any OpenAI-compatible base URL, with a per-slot fallback chain that survives provider outages.
- Strength: inspectable memory. Memories are plain text rows with embeddings, listed and edited through
/api/memories, and switchable per server and per user. - Caveat: heavy footprint. Postgres with pgvector is mandatory, and the default embedding and cross-encoder models pull PyTorch into the image.
- Caveat: shallow default loop. Outside
/research, a turn is one tool-selection call followed by a fixed sequence of tools. Iterative planning and MCP tools only happen in research mode. - Caveat: little action gating. Searches and sandboxed code run without confirmation. The operator is the one place that can pause for the user, and it does so only when the model asks.
- Caveat: SaaS shape. Subscriptions, price tiers and per-plan rate limits are part of the routing code. With billing off, every user counts as subscribed (adapters), but the free/paid model slots and the hosted-service defaults still shape the admin setup.
Sources: code at ae229ca, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the Open-source personal assistants questions
Each answer was drafted by a code-reading agent at commit ae229ca. Its citations were checked mechanically. Compare with the other open-source personal assistants →
How is the assistant architected?
answeredKhoj is a Python application with a FastAPI backend and a Next.js (React) frontend served together in a single Docker image. The entry point is main.py:run() which initializes Django (for ORM, migrations, admin), sets up FastAPI with CORS middleware, mounts the Django app at /server, and starts a Uvicorn (dev) or Gunicorn (prod) server.
Agent loop and runtime. The core agent loop lives in api_chat.py:event_generator() — a large async generator that processes each user request step by step. It first detects the conversation command from the user's query (e.g., /notes, /online, /code, /research, /operator) via helpers.py:get_conversation_command(). For the "default" command, an LLM call determines which tools to use (aget_data_sources_and_output_format). Tools are then executed sequentially in order: (1) research iterations, (2) document/note search (semantic vector search over user content), (3) online search, (4) webpage reading, (5) code execution, (6) operator/computer-use. Each tool yields status events back to the client via SSE (HTTP) or a buffered WebSocket protocol. After all tools run, agenerate_chat_response() dispatches to the appropriate LLM provider (OpenAI, Anthropic, or Gemini), streams the response back, and saves the conversation turn.
Frontend/backend split. The frontend is a Next.js app (in src/interface/web/) built with Bun. FastAPI serves it as static files and also provides a Jinja2-based web client for the app and a landing page. Django provides the admin panel (via django-unfold) and the database ORM.
Main packages. The khoj/ package contains: routers/ (API endpoints), processor/ (conversation models, content ingestion, operator agents), database/ (Django models and adapters), search_filter/ (date/file/word query filters), search_type/ (text/image search backends), and utils/ (config, helpers, state).
User request to action flow. A request arrives at an HTTP POST or WebSocket endpoint → FastAPI authenticates (Django session, bearer token, or anonymous mode) → event_generator is invoked → command detection → tool execution phase (each tool calls the LLM as needed) → final LLM call generates the natural-language response with all compiled context → conversation is saved to PostgreSQL via save_to_conversation_log().
How are integrations (email, calendar, chat, docs) implemented?
answeredKhoj integrates with several external services, each implemented differently:
Notion. OAuth 2.0 flow handled in routers/notion.py. The user is redirected to Notion's authorization page; the callback exchanges the code for an access token (using Basic Auth with NOTION_OAUTH_CLIENT_ID/SECRET), stores it in the NotionConfig model (Django DB), and triggers configure_content() to index Notion pages. No refresh-token logic — the token is stored as-is.
GitHub. processor/content/github/github_to_entries.py uses a Personal Access Token (PAT) stored in GithubConfig to clone repo file trees via the GitHub REST API (e.g., api.github.com/repos/{owner}/{name}/contents/). Supports multiple repos per user. Rate-limit aware: waits on X-RateLimit-Reset headers.
Email. routers/email.py uses Resend (via the resend Python SDK) to send magic-link login emails, welcome emails, and feedback forms. Configured via RESEND_API_KEY env var. Falls back silently if not configured.
Online search and web scraping. processor/tools/online_search.py supports six search/read providers: Google Search API, SerperDev, SearXNG, Exa, Firecrawl, and Olostep. Each is configured via env vars (GOOGLE_SEARCH_API_KEY, SERPER_DEV_API_KEY, KHOJ_SEARXNG_URL, etc.) and the active scraper is a model-priority chain in WebScraper/ServerChatSettings. Direct HTTP fetching is used via aiohttp and BeautifulSoup/markdownify for webpages.
MCP (Model Context Protocol). processor/tools/mcp.py implements a generic MCP client that connects to external servers via stdio (for .py/.js scripts) or SSE (for HTTP URLs). MCP servers are stored in the McpServer Django model. The MCPClient class lists available tools and calls them by name.
Phone (Twilio). Optional phone integration via routers/api_phone.py / routers/twilio.py.
Audio. Text-to-speech via ElevenLabs (processor/speech/text_to_speech.py). Speech-to-text via OpenAI Whisper.
Image generation. Via OpenAI DALL-E, Replicate, or Google Imagen (TextToImageModelConfig).
All integrations are on-demand: triggered by the user query, not synced in the background. Content indexing (Notion, GitHub, local files) is triggered manually or via the APScheduler-based periodic indexer.
How is memory and user context stored and retrieved?
answeredKhoj has three memory layers, all stored in PostgreSQL with pgvector for vector embeddings:
1. Conversation history. Chat sessions are stored as JSON blobs in Conversation.conversation_log (a Django JSONField). Each user message and assistant response is saved via save_to_conversation_log() in processor/conversation/utils.py. Messages include references, online results, code context, operator trajectories, images, and thoughts — everything needed to reconstruct the agent's full state. Conversations are keyed per user and optionally per client application. The generate_chatml_messages_with_context() function reconstructs the prompt by reading the last N messages, truncating oldest ones to fit the model's context window.
2. User content (notes/docs). User-uploaded files (markdown, org-mode, plaintext, PDF, images, DOCX) are ingested by content processors in processor/content/. Each file is chunked into Entry objects with a vector embedding (pgvector VectorField), raw text, compiled text, heading, and file metadata. When the user asks /notes, semantic search is run against these entries using cosine distance via EntryAdapters.
3. Long-term memories (UserMemory). UserMemory is a vector-indexed store derived automatically from conversations. After every chat turn, ai_update_memories() (in helpers.py) calls the LLM to extract new facts from the conversation and saves them via UserMemoryAdapters.save_memory(). Two retrieval modes exist: pull_memories() (recent, time-windowed — "medium term") and search_memories() (semantic vector search — "long term"). Memories are injected into the chat prompt as a special <retrieved_memories> block right before the user's query.
Memory configuration. Server admins can disable memory entirely (ServerChatSettings.memory_mode = DISABLED) or set defaults on/off. Users can toggle it in UserConversationConfig.enable_memory.
Memories are accessible and editable via the /api/memories REST endpoints (list, update, delete).
How are actions on the user's behalf gated?
answeredKhoj has limited human-in-the-loop gating — actions are generally executed autonomously once triggered, with a few specific safeguards:
Operator/computer-use safety. The most powerful actions (clicking, typing, scrolling, file editing, running terminal commands) are executed inside the operator loop (processor/operator/__init__.py). The loop has a key safety valve: it can yield a RequestUserAction action type (operator_actions.py:RequestUserAction) which pauses the loop and asks the user for input via the status event stream. When this happens, the environment is left open (not closed) to wait for user response. The loop also has a hard iteration limit (KHOJ_OPERATOR_ITERATIONS, default 100) and responds to client disconnect cancellation events.
Rate limiting. Three tiers: ApiUserRateLimiter (per-minute and per-day for chat), ApiImageRateLimiter (max images and size), and ConversationCommandRateLimiter (per-command, with trial=20 / subscribed=75 limits). These are enforced as FastAPI dependency injections on the chat endpoint.
Subscription gating. ConversationAdapters.get_chat_model() splits users by subscription: subscribed users can use any ChatModel including paid-tier models; free/trial users are restricted to price_tier = FREE models. Server admins configure which models are free/paid via the Django admin.
Process locks. The ProcessLock model with named operations (index_content, schedule_leader, apply_migrations) provides distributed mutual exclusion via PostgreSQL. Used to ensure only one worker indexes content at a time and only one worker acts as scheduler leader.
No generic approval flow. There is no general "confirm this action before execution" flow for note searches, web searches, or code execution — those run automatically. The operator's RequestUserAction is the only mechanism that explicitly pauses for user consent. There is no persistent audit trail of actions performed beyond conversation history (which records what the assistant chose to do).
How are LLM providers selected and configured?
answeredKhoj supports three LLM providers with a model-agnostic routing layer:
Supported providers. OpenAI (ChatModel.ModelType.OPENAI), Anthropic (ANTHROPIC), and Google Gemini (GOOGLE). Each ChatModel row stores the model name, a foreign key to AiModelApi (which holds api_key and optional api_base_url), a model_type, vision_enabled flag, price_tier, and max_prompt_size. The AiModelApi abstraction allows any OpenAI-compatible endpoint to serve as a provider — the base URL is fully configurable, enabling local models via Ollama, VLLM, LMStudio, etc.
Config surface. Server admins configure models via the Django admin panel (backed by ServerChatSettings). Models are slotted into positions: chat_default (free), chat_advanced (paid), plus fast/deep variants for free and paid users. Slots are resolved per-request via aget_chat_model_slot(). Users can override their model via UserConversationConfig.setting. Model fallback is chained: send_message_to_model_wrapper() tries the primary model first, then iterates through all fallbacks (from all ServerChatSettings entries ordered by priority), retrying on rate-limit/API/timeout errors. Non-retryable errors (auth, invalid model) propagate immediately.
Tool-calling/structured output. The utils.py for OpenAI (processor/conversation/openai/utils.py) detects StructuredOutputSupport level per model: TOOL (native function/tool calling), SCHEMA (JSON Schema in response_format), OBJECT (basic json_object), or NONE. For Anthropic, tool definitions are passed in the Anthropic tool-use format. Structured output schemas use the OpenAI response_format with json_schema and Anthropic's tool-use respectively. For non-schema-capable models, clean_json() strips markdown fences and falls back to string parsing.
Local-model support. Any OpenAI-compatible API endpoint works: set KHOJ_DEFAULT_CHAT_MODEL and OPENAI_BASE_URL (e.g., to http://localhost:11434/v1/ for Ollama). The is_local_api() helper detects localhost URLs and skips HTTPS upgrades.
Other model types: Text-to-image (OpenAI DALL-E, Replicate, Google Imagen), speech-to-text (OpenAI Whisper), text-to-speech (ElevenLabs), and embedding models (sentence-transformers from HuggingFace, configurable via SearchModelConfig).
How is it deployed and self-hosted?
answeredKhoj is designed for Docker-based self-hosting with a single docker-compose.yml that orchestrates five services:
Required dependencies. PostgreSQL with pgvector (pgvector/pgvector:pg15) is the sole hard database dependency — it stores all data: users, conversations, vector embeddings, content entries, and scheduler state. A SearxNG container provides privacy-respecting meta web search. A Terrarium sandbox container (ghcr.io/khoj-ai/terrarium) provides isolated Python code execution. An optional Docker-based computer container (ghcr.io/khoj-ai/khoj-computer) enables the operator/computer-use feature via a VNC-accessible desktop environment.
Docker/one-click paths. The Dockerfile builds a single image containing both the Python server and the compiled Next.js frontend. It installs system deps (swig, musl, libsqlite), Pip-installs Python packages with CPU-only PyTorch, and runs collectstatic. The prod.Dockerfile does the same but uses gunicorn instead of Uvicorn and includes the [prod] extras (Stripe, Twilio, S3). The standard docker-compose up works with minimal env vars (only Postgres credentials and a Django secret key are required). The server automatically runs Django migrations on startup via main.py.
Required external accounts. To actually use the assistant, you need at least one LLM provider API key: OpenAI (OPENAI_API_KEY), Anthropic (ANTHROPIC_API_KEY), or Google Gemini (GEMINI_API_KEY). For web search, you need a Google Search API key (GOOGLE_SEARCH_API_KEY + engine ID) or a SerperDev/Firecrawl/Exa key. Email login requires a Resend API key. Production deployments should configure KHOJ_DOMAIN and a strong KHOJ_DJANGO_SECRET_KEY.
Scheduler. APScheduler (with Django job store) handles periodic tasks: content re-indexing (every 22-25 hours), telemetry upload (every 2 min), stale rate-limit cleanup (every 31 min). A leader-election mechanism via ProcessLock ensures only one worker executes scheduled tasks in multi-worker deployments.
Storage. Image assets can optionally be stored in S3 (configured via AWS_ACCESS_KEY/AWS_SECRET_KEY). Logs go to stdout (Rich) and a log file. No Redis, no message queue — the scheduler piggybacks on the web worker.
schedule library, polled every 60 s by poll_task_scheduler; APScheduler with the Django job store is used for user automations, and the ProcessLock leader election gates that scheduler.