LLMs Technical Reviews

How is retrieval (RAG) implemented?

Embedding models; vector store; top-k; how retrieved code reaches the prompt; or agentic file reading instead of RAG.

Verdict

This is the core split in the category.

deepwiki-open is classic RAG. It uses adalflow with a FAISSRetriever (top_k: 20) over embeddings from OpenAI text-embedding-3-small (256 dims) by default, or Google, Ollama or Bedrock through DEEPWIKI_EMBEDDER_TYPE. The index is a pickle on disk. In research_chat(), the retrieved chunks are grouped by file, labelled with [lines a-b], and placed between <START_OF_CONTEXT> tags. The query is the last user message, which for wiki pages is the whole page prompt. If that message is over 7,500 tokens, retrieval is skipped completely.

OpenDeepWiki has no embeddings and no vector store. The agent reads code through tool calls: ListFiles, Grep and ReadFile on GitTool, plus ReadDoc over pages it has already generated when it answers chat. During generation a budget wrapper limits this to 6 calls per page (4 for the catalog) and 240 lines per read.

RAG is cheaper per page and finds scattered snippets that match a question, but the model never sees whole files. Agentic reading gives real file context and precise source lists, but it depends on the model searching well within a small budget. For Q&A over very large codebases, deepwiki-open’s index finds more. For pages that explain one subsystem in depth, OpenDeepWiki’s approach usually reads more relevant code.

Per-project answers

AsyncFuncAI/deepwiki-open

answered

Retrieval is implemented via a FAISS-based RAG pipeline using the adalflow framework. The core class is RAG in api/rag/rag.py (line 134), which wraps a FAISSRetriever (line 290) configured with top_k: 20 (from embedder.json). Documents are embedded using configurable embedders: OpenAI text-embedding-3-small (256 dims) by default, Google gemini-embedding-001, Ollama nomic-embed-text, or Bedrock amazon.titan-embed-text-v2:0 — selected via the DEEPWIKI_EMBEDDER_TYPE env var (config.py line 65).

The retrieval pipeline in research_chat() (api/services/research.py, line 60) first prepares the retriever index (loading from a .pkl pickle file or creating it fresh), then calls rag.acall(query) which invokes the FAISS retriever. Retrieved documents are grouped by file path and each chunk is annotated with its line range (line 178-184). The grouped context is assembled into a structured prompt that includes file-path headers and line-numbered chunks inside <START_OF_CONTEXT> / <END_OF_CONTEXT> tags (line 146-192).

The context is then injected into prompt_builder() (api/chat/_prompts.py, line 6), which builds a final prompt combining the system prompt, conversation history, context block, and user query. The response streamer (ChatStreamer in _stream.py) sends the full prompt to the LLM and streams back chunks. If the input exceeds ~7500 tokens, RAG context is omitted and a fallback note is inserted instead (line 27-29 of _prompts.py). The Memory class in rag.py (line 79) accumulates conversation turns for multi-turn dialogue.

AIDotNet/OpenDeepWiki

answered

There is no traditional RAG pipeline — no vector embeddings, no vector database, and no top-k similarity search. Instead, OpenDeepWiki uses agentic file and document reading via AI tool calls, implemented in two key classes.

Pre-generated document retrieval. The DocReadTool (src/OpenDeepWiki/Agents/Tools/DocReadTool.cs:14) exposes ReadDocumentAsync(path) and ListDocumentsAsync() to the AI chat agent. It looks up the repository, branch, and language in EF Core (DocCatalog + DocFile entities) and returns cached Markdown content. Each document record stores a SourceFiles JSON array (the source code files that were analyzed to generate it). The ChatDocReaderTool (src/OpenDeepWiki/Agents/Tools/ChatDocReaderTool.cs:15) provides a more granular ReadAsync(path, startLine, endLine) capped at 200 lines per request, with dynamic tool descriptions that embed the available-document catalog.

Source-code retrieval. The GitTool (src/OpenDeepWiki/Agents/Tools/GitTool.cs:13) exposes three AI tools: ReadFile(path, offset, limit) (line 180), ListFiles(glob, maxResults) (line 493), and Grep(pattern, glob, ...) (line 283). These let the agent read arbitrary source files (respecting .gitignore and hidden-path filtering), glob for project structure, and regex-search across files. Lines are capped at 2000 characters; reads default to 2000 lines; grep results default to 50 with 2 context lines. Binary files are skipped by extension.

How it reaches the prompt. In ChatAssistantService.StreamChatAsync() (src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:372), tools are constructed at request time: a GitTool is initialized from the cloned repository path on disk, a ChatDocReaderTool loads the document catalog from the database, and optionally MCP/skill tools. The system prompt (line 963–1138) instructs the agent to first ReadDoc for documentation, check the sourceFiles field, then use ReadFile/Grep for deeper source code investigation. The DocumentSourceToolBudget (src/OpenDeepWiki/Agents/Tools/DocumentSourceToolBudget.cs) caps how many source-discovery tools (ListFiles/Grep/ReadFile) the agent can use per page.

← How is a repository ingested and chunked? · How is the wiki structure (table of contents) determined? →