How is repository context gathered and kept within the context window?
Repo maps, search/grep tools, embeddings, file-reading strategy; summarisation or compaction of long sessions.
Verdict
Aider is the only coding agent that builds a global view of the repository up front. The others let the model search on demand and differ mainly in how they compact long sessions. None of the twelve uses embeddings.
Precomputed repo map. Aider ranks tree-sitter definitions and references with personalised PageRank. The map fits a budget of 1/8 of the input window, clamped to 1K–4K tokens.
On-demand tools plus compaction.
- Cline: reads up to 2,000 lines, caps each tool result at 8,000 characters, and compacts at 90% of the window down to 70%.
- Codewhale: prunes old tool output before any LLM summary and keeps the last user round word for word.
- DeepSeek-Reasonix: a tree-sitter
code_indexgives outlines without a language server. At 85% it writes a structured digest of at most 16K tokens, keeping the prompt prefix cache-stable. - Qwen Code: replaces old tool outputs with
[Old tool result content cleared]and runs background auto-memory under~/.qwen/memories/. - OpenCode: prunes tool outputs while protecting the last 40K tokens, then hands off to a dedicated
compactionagent. - PI-Desktop: a model-written summary by default. It throws rather than send an oversized request.
- Codex: no dedicated read or search tools. The model reads with shell commands, each going through approval and the sandbox.
- Open Interpreter: Codex’s compaction, keeping up to 20,000 tokens of recent user messages beside the summary.
Per-request selection in the app builders.
- bolt.diy: two extra LLM calls per request, a chat summary and a file selection. Only the last three messages are kept, and the five-file limit is a prompt instruction, not code.
- Open Lovable: a structured search plan drives a local grep. Context files are cut to 2,000 characters, and nothing is summarised.
- VibeSDK: Think trims history but keeps tool pairs around
deploy_spaceandcommit.
Pick: Aider for a cheap, global map of a large repository. Pick: DeepSeek-Reasonix or Codewhale for long sessions where prompt-cache hits drive cost. Pick: bolt.diy if you want context chosen per request without hand-picking files.
More projects in this category are being researched.
Per-project answers
anomalyco/opencode
answeredContext Management
OpenCode gathers repository context through dedicated tools rather than embedding a full repo map directly in the prompt. The read tool (tool/read.ts) reads file contents with optional line ranges. The grep tool (tool/grep.ts) does regex content search. The glob tool (tool/glob.ts) does pattern-based file listing. The list tool provides directory listings. There is also an MCP resource system that can fetch context from connected servers — list_mcp_resources, list_mcp_resource_templates, and read_mcp_resource (session/tools.ts lines 139-386).
System prompts are assembled per step in runLoop (prompt.ts lines 1257-1270): environment info, agent instructions from AGENTS.md-style files (instruction.ts), MCP instructions, and skills. The environment prompt in session/system.ts (lines 69-80) includes the working directory, workspace root, OS, current date/time, and the model ID. Per-model prompt variants exist for Anthropic, GPT, Gemini, Kimi, Meta, Codex, and a generic default (session/system.ts lines 28-51, session/prompt/ directory).
Context window overflow is detected by isOverflow in session/overflow.ts (lines 8-34), which checks if total tokens exceed the model's context limit minus a reserved buffer (default 20k). The SessionCompaction service (session/compaction.ts) handles overflow via pruning (lines 273-317) — walking backwards through recent tool calls, erasing the output text of older ones (marking compacted timestamp) to free space while keeping at least PRUNE_PROTECT (40k) tokens protected — and compaction (lines 319-557), which uses a dedicated "compaction" agent to summarize the conversation history. The compaction strategy selects a tail_start_id based on preserveRecentBudget (lines 115-119, default 2k-15k tokens reserved for recent turns), and the compaction agent generates a summary that prepends the next model call. If auto-compaction is enabled, after compaction a synthetic "Continue" message is injected (line 519-548).
There is no embedding-based RAG or vector search in the codebase; context retrieval relies entirely on tool calls (the model decides what to grep/read/glob) and compaction-based summarization of long sessions.
openai/codex
answeredRepository context is built through a layered World State system in core/src/context/world_state/mod.rs:64. Each turn constructs a WorldState by merging context fragments from many sources: base instructions, model instructions, AGENTS.md instructions (loaded in core/src/agents_md.rs:58 by walking up from the project root to find files named AGENTS.md), environment configurations, plugin/skill instructions, and permission profiles. These fragments are placed as Prefix, Standalone, or Mergeable items (line 70–75) into the model's input. File search uses a fuzzy-file-search tool and a ripgrep-based content search tool (exposed through the unified exec tool surface). For context-window management, Codex CLI implements auto-compaction in core/src/compact.rs:1. When the token count exceeds the model's limit — tracked via ContextWindowTokenStatus in core/src/session/context_window.rs:8 — run_auto_compact() (called from turn.rs at lines 625, 734, 785) sends the conversation history to the model with a summarization prompt (SUMMARIZATION_PROMPT from prompts). The two main compaction strategies are inline auto-compact (line 625) and remote auto-compact v2 (imported from compact_remote_v2). Session history is stored via the thread store and cloned for each model request, with compaction producing CompactedHistoryMetadata that replaces verbose history. The WorldState also tracks per-section context diffs and token budgets, emitting context-window guidance fragments when approaching limits.
cline/cline
answeredContext is managed by the MessageBuilder (message-builder.ts) in @cline/core. It walks the conversation history and produces provider-ready message payloads, applying two compaction strategies: tool-result truncation (each tool result is capped at DEFAULT_MAX_TOOL_RESULT_CHARS = 8,000 chars) and outdated-file-content rewriting (re-read files replace stale content with [outdated - see the latest file content]). There is an aggregate text budget of 6 MB (DEFAULT_MAX_TOTAL_TEXT_BYTES) to prevent oversized requests. On context-window overflow, the runtime attempts recovery by compacting the conversation and retrying once; if that fails, it terminates with a specialized error message (agent-runtime.ts:106-133). The file-reading strategy uses line-range reads (file-read.ts) with limits of 10 MB file size and 50,000 un-ranged lines, and images are read as base64 inline content. Search/grep is a built-in tool backed by ripgrep via the search executor (search.ts). Checkpoints snapshot the workspace via git stash with a private ref (checkpoint-hooks.ts, checkpoint-restore.ts). The orchestrator (session-runtime-orchestrator.ts) ties these together per-session with a ConversationStore for the message transcript, a ToolResultCache for caching tool outputs, and a MistakeTracker/LoopDetectionTracker for safety.
extensions/context/compaction.ts: it triggers at 90% of the input window, defaults to the agentic (LLM summary) strategy with a basic truncation fallback, targets 70% and keeps the last 20K tokens. Checkpoint snapshots use git stash create, not git stash.openinterpreter/openinterpreter
answeredContext is managed through compaction, AGENTS.md discovery, and per-turn assembly. The StepContext captures world state, environment snapshot, tool router, MCP binding, and model info. History is cloned from the session and filtered by model input modalities (turn.rs:513-519). Pre-turn compaction via run_pre_sampling_compact() (turn.rs:183-221) calls run_auto_compact() (compact.rs:1446-1504) which delegates to local summarization or remote compaction v2 depending on the provider's RemoteCompactionSupport. Mid-turn compaction triggers on context window exceeded (turn.rs:614-641); post-turn compaction can also run (line 706-735). The summary is capped at 20,000 tokens (compact.rs:62). Repository context comes from AGENTS.md files (core/src/agents_md.rs:1-18): walking from project root (found via .git markers) to cwd, concatenating into user instructions. Fallback: AGENTS.override.md. File-search tools (core/src/tools/handlers/tool_search.rs:28-33) provide BM25 search over tools. There are no vector embeddings or repo maps -- context is prompt concatenation and compaction.
COMPACT_USER_MESSAGE_MAX_TOKENS, compact.rs L62) caps the recent user messages retained next to the summary in the rebuilt history, not the summary length.Aider-AI/aider
answeredRepo map via tree-sitter tags. The RepoMap class (aider/repomap.py:42) scans all tracked files in the git repo with tree-sitter to extract symbol definitions (functions, classes, etc.) as Tag objects. The method get_ranked_tags_map() (repomap.py:576-627) ranks tags by relevance — files already in the chat rank higher, then files with mentioned identifiers or filenames. A binary search (repomap.py:676-700) iteratively includes more tags until the map fills the allocated token budget (max_map_tokens, default 1024). The result is a compact tree showing file names and their key symbols. The map is cached in an LRU dict and also in a diskcache-backed SQLite DB at .aider.tags.cache.v3/ for cross-session reuse.
File content in messages. The Coder.format_chat_chunks() method (base_coder.py:1226) assembles messages in this order: system prompt, examples, repo map (as user/assistant pair), readonly file contents, done/archived messages, editable file contents, current conversation, reminder. File contents are read from disk via get_files_content() and inserted as literal text. Image files are base64-encoded and sent as image_url content blocks (get_images_message() at base_coder.py:817-850).
File-add by mention. When the LLM mentions a file not yet in the chat, check_for_file_mentions() (base_coder.py:1761-1781) detects the mention and asks the user to add it. This is a form of dynamic context expansion.
Summarisation of long sessions. The ChatSummary class (aider/history.py) runs in a background thread (summarize_start() at base_coder.py:1002-1012). When done_messages exceed max_chat_history_tokens, the summarizer trims older messages by splitting at the halfway point and using the weak model to compress the head segment via summarize_all() (history.py:98), keeping the more recent tail intact. Depth is capped at 3 recursive splits.
Prompt caching. When enabled, ChatChunks.add_cache_control_headers() (aider/coders/chat_chunks.py:28-55) adds cache_control: ephemeral markers to the last message of cacheable segments (examples/system, repo/readonly, chat files). A background thread (warm_cache() at base_coder.py:1340-1394) periodically sends 1-token pings to keep the cache alive.
codewhale-hq/Codewhale
answeredCodewhale gathers repository context through explicit tool calls rather than automated crawling or embedding-based retrieval. There are no embeddings or vector indexes — the model reads files it chooses.
File reading: The read/ReadTool and legacy ReadFileTool (tools/file.rs:1092) stream file contents with: small-file fast-path (whole file in one call when ≤4KB and ≤500 lines), line-range windows (start_line/max_lines), byte-size awareness with truncation footers naming the next offset, content hashing (SHA-256) reported in responses so the model can validate later edits. Binary files route to OCR (images) or PDF extraction.
Search tools: search_content (Grep) and search_name (FileSearch) provide pattern-based discovery (tools/search.rs, tools/file_search.rs). A deferred tool_search tool (tool_catalog.rs:35) with regex/BM25 matching allows discovery of tools and files by name or content. Web search is available via web_search.
System prompt assembly: The system prompt (prompts/text.rs:44, BASE_PROMPT) is composed from layers: constitution (binding core), personality overlay (currently CALM), approval-policy overlays, memory block (user memories from runtime/src/native_memory.rs), skills index, MCP server instructions, mode-specific instructions (Plan/Act/Operate), and project instruction files (CLAUDE.md/AGENTS.md). The session-pinned prefix is kept KV-cache stable — volatile facts (steers, completions, diagnostics) are appended as user-role messages, never spliced into the frozen prefix.
Compaction for long sessions (compaction.rs:1): When context pressure exceeds a model-specific token threshold (80% of context window), compaction triggers. The strategy is: first prune old tool results from the tail (cheapest, no LLM call), then if still over threshold, run an LLM summarization pass that compresses older user+assistant turns while keeping the last user round verbatim (the survival contract in survival_contract.rs, enforced by last_round.rs). Recent plain user messages are retained up to a configurable token budget (retained_user_message_tokens). The pre-compaction history is persisted as an artifact before any mutation. Context pressure warnings (warning at 70%, critical at 85%) are emitted mid-turn via context_pressure_message.
Memory: User memories (tools/native_memory.rs, tools/remember.rs) provide a persistent capture path — the model can ask to remember things, which are stored on disk and prepended to the system prompt on subsequent sessions.
esengine/DeepSeek-Reasonix
answeredRepository context is gathered through compile-time-builtin tools, not embeddings or a repo-map pass. The primary tools are: read_file (internal/tools/builtin/readfile.go:52) reads files with offset/limit paging and reports total-length/pagination trailer; grep (internal/tools/builtin/grep.go:30) runs ripgrep underneath, capped at 200 results and 300s timeout; glob finds files by glob pattern; code_index (internal/tools/builtin/codeindex.go:51) provides outline/search for Go/JS/Python/Rust/TypeScript via tree-sitter AST parsing without a language server. The model selectively invokes these on demand — there is no automatic repository indexing. Context window management is handled by windowState (internal/runtime/agent/context_window.go:13) with a compaction system (internal/runtime/agent/compact.go:20). At 85% capacity (defaultCompactRatio), the agent installs a summary checkpoint: earlier conversation is replaced by a structured brief under headings (Standing facts, Goal, Decisions, Files, Commands, Errors, Pending) generated by the model itself with a low-effort summarizer prompt (compact.go:57). The recent tail (at least 10% of window, 32K–96K tokens) and a small number of verbatim user turns are preserved. Compaction is one-shot per maintenance boundary; the model sees a <compaction-summary>...</compaction-summary> block. The workspace scan (workspace_scan.go:22) walks the workspace tree (skipping VCS dirs, symlinks) in parallel (16 readers) up to 50K files, used for delivery-verification scans rather than continuous context. There is no embedding-based retrieval or RAG system.
firecrawl/open-lovable
answeredRepository context is gathered through a FileManifest system (types/file-manifest.ts:44-51) that records every file's full content, exports, imports, component dependencies, and last-modified time. For edit requests, context flows through a multi-stage pipeline. First, analyzeEditIntent() in lib/edit-intent-analyzer.ts:6 uses regex pattern-matching against the user prompt to classify the edit type and resolve target files. A more advanced path calls the /api/analyze-edit-intent endpoint (app/api/analyze-edit-intent/route.ts:126-168) which uses generateObject() with a Zod schema to let the AI produce a search plan with search terms, regex patterns, and file-type filters. Then executeSearchPlan() (lib/file-search-executor.ts:59-133) performs line-by-line grep over file contents. Finally formatFilesForAI() (lib/context-selector.ts:308-347) builds the system prompt: primary files shown in full, context files truncated to 2000 chars. The conversation state (types/conversation.ts:45-49) is tracked in global variables and includes message history, edit records, and project evolution. Context window management is aggressive: messages are capped at 20 (trimmed to 15), edits at 10 (trimmed to 8), conversation context truncated to 2000 chars (generate-ai-code-stream/route.ts:130-575). There are no embeddings, no vector search, and no summarisation or compaction of long sessions beyond simple splicing. Web-scraped content (from the Firecrawl API) is included as-is with a 1000-char preview.
QwenLM/qwen-code
answeredRepository context is gathered through several mechanisms. getDirectoryContextString() in packages/core/src/core/environmentContext.ts (lines 53–83) reads folder structures via getFolderStructure() for all workspace directories and injects them as a system prompt section. The getEnvironmentContext() function (lines 92–100) assembles the current date, OS platform, and directory tree into the startup context. IDE context blocks are generated per‑turn by LlmClient.sendMessageStream() and wrapped as <system-reminder> tags (lines 4429–4443 of client.ts).
Auto-Memory is the primary persistent-memory system. MemoryManager (packages/core/src/memory/manager.ts) runs background extract/dream/recall/forget cycles, building a tree of markdown memory files under .claude/memory/. The auto-memory prompt is injected into the system prompt before each turn. Microcompaction (packages/core/src/services/microcompaction/microcompact.ts) replaces old tool outputs (Read, Grep, Shell, WebFetch, etc.) with placeholder text ([Old tool result content cleared]), keeping the conversation within context limits. Chat compression (packages/core/src/services/chatCompressionService.ts) runs a separate model side-query to summarize the conversation when token thresholds are breached, then replaces the compacted history with a summary. File content is read on-demand via dedicated tools (Read, Grep, Glob) — there is no automatic full‑repo embedding or vector index.
~/.qwen/memories/ (user scope) and a per-project auto-memory root, not .claude/memory/ (packages/core/src/memory/paths.ts).stackblitz-labs/bolt.diy
answeredContext gathering is a multi-phase server-side process triggered on every request when contextOptimization is enabled. First, createSummary (app/lib/.server/llm/create-summary.ts) sends all processed messages to the LLM with a structured prompt requesting a summary in a defined template format (project overview, decisions, implementation status, etc.). The summary is stored as a chatSummary annotation on the data stream. Next, selectContext (app/lib/.server/llm/select-context.ts) retrieves all file paths from the project (filtered by an ignore-list similar to .gitignore via the ignore package), reads the current context buffer from annotations, and asks the LLM which files are relevant by outputting <updateContextBuffer> XML with <includeFile> / <excludeFile> tags. Only 5 files are selected at a time (line 170). The file content is read from the in-memory FileMap (populated by FilesStore from the WebContainer filesystem) and formatted as boltAction type="file" entries inside a boltArtifact wrapper via createFilesContext in app/lib/.server/llm/utils.ts. For long sessions, messageSliceId is set to messages.length - 3 (api.chat.ts line 106-107), keeping only the last 3 messages, while the summary carries forward the earlier history. The summary and code context are injected into the system prompt before the main LLM call. The system prompt builder getSystemPrompt includes environment constraints, database instructions, artifact formatting rules, and design guidance — all composed into a single system message.
vastsa/PI-Desktop
answeredContext is gathered from session history and managed via compaction. The SessionRuntime maintains a fullEntries array of Entry objects (message, compaction, branch_summary, custom) — the append-only transcript. On each turn, buildSessionContext() in session-context.ts:94-105 projects these entries into { messages: AgentMessage[] } for the model, slicing from the latest compaction entry forward (buildContextEntries(), session-context.ts:43-51).
Compaction — When the context budget is near its limit (automaticCompactionThresholdFor()), prepareNextTurn() (runtime.ts:6560-6609) triggers compaction before the next provider request. This runs a pi-based summary: earlier conversation is summarized into a compact text block, the retained tail (latest user messages, pending tool results) is preserved, and a checkpoint is persisted. The strategy defaults to summary (spends a model request to generate a structured summary) but can be set to fresh_window via env var (runtime.ts:887-894). Fallback paths exist when summary generation fails (COMPACTION_FALLBACK_MARKER, runtime.ts:701-704).
Repo maps and search — The system prompt sections include a runtime section (mode prompt, tool list, shell guidance, scratch dir) and tool activation state (system-transcript.ts:36-65). The model calls tools like Grep (regex search), Glob (file pattern matching), and Read (bounded file windows with line numbers) to gather context at runtime — there is no offline embedding or vector index.
Long sessions — Context compaction is the primary mechanism for keeping sessions within the window. The compaction entry replaces old messages with a summary plus a retained tail of recent messages. Delegation helps too: subagent runs produce a compact report rather than their full transcript. The context_budget system section reminds the model of remaining tokens (runtime.ts:938-958).
cloudflare/vibesdk
answeredRepository context is gathered and managed through several mechanisms:
Codebase reading tools: In the Think loop, workspace tools include read, list, find, and grep from @cloudflare/think/tools/workspace which operate on the SpaceDO workspace. The ThinkAgent has no separate glob tool beyond the generic workspace tools, but the system prompt instructs the model to use search tools extensively. The Gemini system prompt (worker/agents/think/prompts/gemini.txt) explicitly tells the agent: "Use grep and glob search tools extensively (in parallel if independent) to understand file structures, existing code patterns, and conventions. Use read to understand context and validate any assumptions."
Context selection: ThinkAgent.beforeTurn() calls selectThinkContextMessages() in worker/agents/think/context-selector.ts to trim the message history before each turn. It preserves all tool-result pairs for tool names matching a protected list (deploy_space, commit, ask_questions), keeps a minimum of RECENT_MESSAGE_COUNT (5) recent messages, and uses a diff-based strategy to preserve complete tool-call/tool-result pairs around named tools rather than naive truncation.
Conversation compactification: The older Agentic/Phasic behaviors use worker/agents/utils/conversationCompactifier.ts which triggers summarization at 40 conversation turns or 100k estimated tokens. It preserves the last 10 messages uncompacted, calls an LLM to summarize older history into a concise context block wrapped in <system_context> tags, and replaces compacted messages with the summary.
File-reading strategy: The Think model is instructed to read files to understand context before making changes. The workspace operations layer reads and writes through SpaceDO RPC calls — not local filesystem — so file I/O goes through Durable Object stubs. The codebaseContext.ts utility provides a CodebaseContextBuilder for the Agentic behavior that can scan directories, collect file contents, and build structured context blocks for the LLM.
VibeSDK does not use embeddings or vector search for repository context. It relies entirely on LLM-driven search (grep/find within workspace tools) and conversation history management.
← How is the agent loop implemented? · How are code edits applied? →