# How is retrieval (RAG) implemented?

> Open-source DeepWiki — a good answer covers: Embedding models; vector store; top-k; how retrieved code reaches the prompt; or agentic file reading instead of RAG.

Canonical page: https://llms-technical-reviews.com/open-source-deepwiki/q/retrieval/

## Verdict

This is the core split in the category.

[deepwiki-open](/p/deepwiki-open/) is classic RAG. It uses adalflow with a `FAISSRetriever` (`top_k: 20`) over embeddings from OpenAI `text-embedding-3-small` (256 dims) by default, or Google, Ollama or Bedrock through `DEEPWIKI_EMBEDDER_TYPE`. The index is a pickle on disk. In `research_chat()`, the retrieved chunks are grouped by file, labelled with `[lines a-b]`, and placed between `<START_OF_CONTEXT>` tags. The query is the last user message, which for wiki pages is the whole page prompt. If that message is over 7,500 tokens, retrieval is skipped completely.

[OpenDeepWiki](/p/opendeepwiki/) has no embeddings and no vector store. The agent reads code through tool calls: `ListFiles`, `Grep` and `ReadFile` on `GitTool`, plus `ReadDoc` over pages it has already generated when it answers chat. During generation a budget wrapper limits this to 6 calls per page (4 for the catalog) and 240 lines per read.

RAG is cheaper per page and finds scattered snippets that match a question, but the model never sees whole files. Agentic reading gives real file context and precise source lists, but it depends on the model searching well within a small budget. For Q&A over very large codebases, deepwiki-open's index finds more. For pages that explain one subsystem in depth, OpenDeepWiki's approach usually reads more relevant code.

## Per-project answers

### AsyncFuncAI/deepwiki-open (answered)

Retrieval is implemented via a FAISS-based RAG pipeline using the `adalflow` framework. The core class is `RAG` in `api/rag/rag.py` (line 134), which wraps a `FAISSRetriever` (line 290) configured with `top_k: 20` (from `embedder.json`). Documents are embedded using configurable embedders: OpenAI `text-embedding-3-small` (256 dims) by default, Google `gemini-embedding-001`, Ollama `nomic-embed-text`, or Bedrock `amazon.titan-embed-text-v2:0` — selected via the `DEEPWIKI_EMBEDDER_TYPE` env var (config.py line 65).

The retrieval pipeline in `research_chat()` (`api/services/research.py`, line 60) first prepares the retriever index (loading from a `.pkl` pickle file or creating it fresh), then calls `rag.acall(query)` which invokes the FAISS retriever. Retrieved documents are grouped by file path and each chunk is annotated with its line range (line 178-184). The grouped context is assembled into a structured prompt that includes file-path headers and line-numbered chunks inside `<START_OF_CONTEXT>` / `<END_OF_CONTEXT>` tags (line 146-192).

The context is then injected into `prompt_builder()` (`api/chat/_prompts.py`, line 6), which builds a final prompt combining the system prompt, conversation history, context block, and user query. The response streamer (`ChatStreamer` in `_stream.py`) sends the full prompt to the LLM and streams back chunks. If the input exceeds ~7500 tokens, RAG context is omitted and a fallback note is inserted instead (line 27-29 of `_prompts.py`). The `Memory` class in `rag.py` (line 79) accumulates conversation turns for multi-turn dialogue.


Citations: [api/rag/pipeline.py:170-208](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/pipeline.py#L170-L208) · [api/services/research.py:60-311](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/research.py#L60-L311) · [api/chat/_prompts.py:6-31](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/chat/_prompts.py#L6-L31) · [api/config/embedder.json:1-41](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config/embedder.json#L1-L41) · [api/config.py:65-65](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config.py#L65-L65)

### AIDotNet/OpenDeepWiki (answered)

There is no traditional RAG pipeline — no vector embeddings, no vector database, and no top-k similarity search. Instead, OpenDeepWiki uses **agentic file and document reading** via AI tool calls, implemented in two key classes.

**Pre-generated document retrieval.** The `DocReadTool` (src/OpenDeepWiki/Agents/Tools/DocReadTool.cs:14) exposes `ReadDocumentAsync(path)` and `ListDocumentsAsync()` to the AI chat agent. It looks up the repository, branch, and language in EF Core (DocCatalog + DocFile entities) and returns cached Markdown content. Each document record stores a `SourceFiles` JSON array (the source code files that were analyzed to generate it). The `ChatDocReaderTool` (src/OpenDeepWiki/Agents/Tools/ChatDocReaderTool.cs:15) provides a more granular `ReadAsync(path, startLine, endLine)` capped at 200 lines per request, with dynamic tool descriptions that embed the available-document catalog.

**Source-code retrieval.** The `GitTool` (src/OpenDeepWiki/Agents/Tools/GitTool.cs:13) exposes three AI tools: `ReadFile(path, offset, limit)` (line 180), `ListFiles(glob, maxResults)` (line 493), and `Grep(pattern, glob, ...)` (line 283). These let the agent read arbitrary source files (respecting `.gitignore` and hidden-path filtering), glob for project structure, and regex-search across files. Lines are capped at 2000 characters; reads default to 2000 lines; grep results default to 50 with 2 context lines. Binary files are skipped by extension.

**How it reaches the prompt.** In `ChatAssistantService.StreamChatAsync()` (src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:372), tools are constructed at request time: a `GitTool` is initialized from the cloned repository path on disk, a `ChatDocReaderTool` loads the document catalog from the database, and optionally MCP/skill tools. The system prompt (line 963–1138) instructs the agent to first `ReadDoc` for documentation, check the `sourceFiles` field, then use `ReadFile`/`Grep` for deeper source code investigation. The `DocumentSourceToolBudget` (src/OpenDeepWiki/Agents/Tools/DocumentSourceToolBudget.cs) caps how many source-discovery tools (ListFiles/Grep/ReadFile) the agent can use per page.


Citations: [src/OpenDeepWiki/Agents/Tools/DocReadTool.cs:14-260](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/Tools/DocReadTool.cs#L14-L260) · [src/OpenDeepWiki/Agents/Tools/ChatDocReaderTool.cs:15-292](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/Tools/ChatDocReaderTool.cs#L15-L292) · [src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:372-1138](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs#L372-L1138)
