# AIDotNet/OpenDeepWiki

> ASP.NET Core service in which tool-calling LLM agents read a repo and write its wiki, mind map and translations into SQLite or PostgreSQL.

- Category: [Open-source DeepWiki](https://llms-technical-reviews.com/open-source-deepwiki/)
- Repository: https://github.com/AIDotNet/OpenDeepWiki (reviewed at commit `d33113c88c56458c202fa6904ec114387319f1f5`, 2026-10-05)
- Stars: 3615 · Language: C# · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/opendeepwiki/

## Overview

OpenDeepWiki is a multi-user wiki platform built on agents instead of RAG. An ASP.NET Core backend (`src/OpenDeepWiki`) imports a repository from Git, a ZIP upload or an approved local directory. It then runs LLM agents (Microsoft.Agents.AI) that call `ListFiles`, `Grep` and `ReadFile` on the working copy and save their output through `WriteCatalog`, `WriteDoc` and `AppendDoc` tools. There are no embeddings and no vector store. Everything is persisted with EF Core in SQLite or PostgreSQL, and a Next.js 16 frontend (`web/`) reads it back.

The scope is much wider than a wiki generator. It has users and roles, organizations, an admin console for AI providers and models, per-branch and per-language documents, background translation, a mind map, incremental updates from `git diff`, a chat assistant with SSE streaming, and an MCP server per repository. It also includes chat-platform providers (QQ and others), which most users will not need.

It suits a team that wants a hosted, shared knowledge base across many repositories and is willing to run a .NET service with a database. It is not a light single-repo tool.

## Architecture

```mermaid
flowchart TD
    API["Submit endpoints"] --> DB[("EF Core: SQLite / PostgreSQL")]
    W["RepositoryProcessingWorker"] -->|"poll 30 s"| DB
    W --> RA["RepositoryAnalyzer.PrepareWorkspaceAsync"]
    W --> SP["RepositoryScanPlanResolver"]
    W --> WG["WikiGenerator"]
    WG --> CAT["GenerateCatalogAsync"]
    WG --> DOCS["GenerateDocumentsAsync"]
    WG --> INC["IncrementalUpdateAsync"]
    CAT --> AG["Agent + tools"]
    DOCS --> AG
    AG --> SRC["DocumentSourceToolBudget / GitTool"]
    AG --> OUT["CatalogTool / DocTool"]
    OUT --> DB
    AG --> AF["AgentFactory"]
    AF --> LLM["OpenAI / Anthropic / DeepSeek / Azure"]
    MM["MindMapWorker"] --> DB
    TW["TranslationWorker"] --> DB
    CHAT["POST /api/v1/chat/stream"] --> DB
    MCP["MCP per repo"] --> DB
```

| Component | Path | Role |
|---|---|---|
| Job loop | [RepositoryProcessingWorker.cs](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryProcessingWorker.cs#L62-L144) | Polls pending repos, takes a lease, runs branches and languages |
| Workspace | `src/OpenDeepWiki/Services/Repositories/RepositoryAnalyzer.cs` | LibGit2Sharp clone/pull with retries, ZIP extract, local copy or link |
| Scan plan | [RepositoryScanPlanResolver.cs](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs#L87-L125) | Sizes the directory tree shown to the catalog agent |
| Generator | [WikiGenerator.cs](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs#L245-L362) | Catalog, documents, incremental update, mind map, translation |
| Source tools | [DocumentSourceToolBudget.cs](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/Tools/DocumentSourceToolBudget.cs#L1-L48) | Budgeted `ReadFile`/`ListFiles`/`Grep` over `GitTool` |
| Output tools | `src/OpenDeepWiki/Agents/Tools/CatalogTool.cs`, `DocTool.cs` | Agent writes catalog JSON and Markdown pages to the DB |
| Model clients | [AgentFactory.cs](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/AgentFactory.cs#L60-L130) | Builds the chat client per request type, with an HTTP handler chain |
| Provider resolver | [AiProviderResolver.cs](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/AI/AiProviderResolver.cs#L137-L227) | Maps DB provider rows to request types and base URLs |
| Prompts | `src/OpenDeepWiki/prompts/*.md` | `catalog-generator`, `content-generator`, `incremental-updater`, `mindmap-generator` |
| MCP | [Program.cs](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Program.cs#L380-L393) | `MapMcp("/api/mcp")` and a per-repo route |

## How a request flows

1. A submit endpoint stores a `Repository` row with status `Pending`. Nothing runs inline.
2. [`RepositoryProcessingWorker`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryProcessingWorker.cs#L18-L60) wakes every 30 s. `DispatchPendingAsync()` takes up to 50 pending or processing repos and asks `IWikiGenerationCoordinator.TryBeginAsync()` for a lease. The lease caps work at `MaxConcurrentGenerations` (default 1) across instances and recovers stale work.
3. [`ProcessBranchAsync()`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryProcessingWorker.cs#L356-L525) calls `PrepareWorkspaceAsync()` with the last processed commit, resolves the scan plan, detects the primary language, and computes changed files if the commit moved.
4. For each branch language, [`ProcessLanguageAsync()`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryProcessingWorker.cs#L530-L599) runs either `IncrementalUpdateAsync()` or `GenerateCatalogAsync()` followed by `GenerateDocumentsAsync()`. Mind maps and translations are left to their own workers.
5. **Catalog.** [`GenerateCatalogAsync()`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs#L245-L362) collects a TOON-serialized directory tree, the README (first 4000 chars), key files and entry points ([#L2411-L2484](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs#L2411-L2484)). It gives the agent 4 source-tool calls and the catalog tools. The run fails if the agent writes no catalog items.
6. **Documents.** [`GenerateDocumentsAsync()`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs#L366-L559) removes duplicate paths and skips leaves that already have content. It generates the first page alone to warm the prompt cache, then runs the rest with `Parallel.ForEachAsync` at `ParallelCount` (default 5). Each page has a 60-minute timeout. A failed page is logged and skipped.
7. **One page.** [`GenerateDocumentContentAsync()`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs#L796-L960) sends the `content-generator` prompt plus a user message with the page title, a File Reference Base URL and budgets: 6 source calls, 4 `AppendDoc` calls. `DocTool` saves the Markdown and records the files the `GitTool` read as `SourceFiles`. `ExecuteAgentWithRetryAsync()` adds a prompt-cache key and retries up to `MaxRetryAttempts` (3).

## Key components

### Agent tools instead of retrieval

Generation never sees embeddings. The agent decides what to read. During generation the tools go through [`DocumentSourceToolBudget`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/Tools/DocumentSourceToolBudget.cs#L1-L48), which caps a read at 240 lines, a listing at 20 files and a grep at 12 results. When the budget is spent it returns `SOURCE_TOOL_BUDGET_REACHED`, and the prompt tells the model to write with what it has. The chat assistant and MCP use the raw `GitTool` with larger defaults (2000 lines per read).

### Scan plan

The catalog agent sees a bounded tree, not the whole repo. [`DecideByRules()`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs#L87-L125) sets the depth (4, 3 above 5000 files, 2 above 20000) and a total-file cap of `min(MaxTotalTreeFiles, 500)`, or 350 above 20000 files. The 800 and 1500 constants are only outer clamps. The plan is saved on the repository so later runs reuse it. Files outside the tree are still readable through `ReadFile` or `Grep`, but the agent has to know to look for them.

### Prompts with renderer rules

`content-generator.md` and the user message in `GenerateDocumentContentAsync()` contain detailed Mermaid rules (quoted labels, `sg_` subgraph IDs, `erDiagram` token rules). There are also two normalizers, `MermaidFlowchartIdNormalizer` and `MermaidMarkdownNormalizer`, in `Services/Wiki`. Pages are written incrementally (`WriteDoc` then `AppendDoc`), so length is not capped by one response.

### Providers

Providers and models are database rows, seeded from `BuiltinProviderPresets.json` (29 presets). [`NormalizeProviderType()` and `ParseRequestType()`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/AI/AiProviderResolver.cs#L137-L208) reduce them to five request types: OpenAI, OpenAIResponses, AzureOpenAI, Anthropic and DeepSeekOpenAI. A `gpt-5` or later model ID goes to the Responses API automatically. Catalog, content and translation can each use a different provider and model (`CatalogModel`, `ContentModel`, `TranslationModel`). `AgentFactory` wraps every call in handlers that fix Gemini `finish_reason` values and carry Gemini thought signatures, with a 300 s HTTP timeout.

### Chat and MCP

`POST /api/v1/chat/stream` streams SSE events (`content`, `thinking`, `tool_call`, `tool_result`, `done`). The agent gets a doc reader over generated pages plus `GitTool` when the clone exists, and sessions are stored in the DB. The global `/api/mcp` endpoint has tools that list repositories, route a question to the best repositories by token scoring, and search or read generated pages (`MCP/McpGlobalTools.cs`). The per-repo endpoint offers [`SearchDoc`](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/MCP/McpRepositoryTools.cs#L22-L107), which finds pages with a SQL `LIKE '%query%'` on title and content and lets an agent summarize them, plus `GetRepoStructure` and `ReadFile`.

## Extending it

- **Providers and models:** add them in the admin console or as preset JSON. Any OpenAI-compatible endpoint works by setting `BaseUrl` with type `openai`.
- **Prompts:** edit the Markdown files in `src/OpenDeepWiki/prompts` (`PromptsDirectory` option).
- **Budgets:** `WikiGeneratorOptions` exposes the tool budgets, parallelism, timeouts, README length and tree limits.
- **Agent tools:** chat accepts configured MCP servers (`McpToolConverter`) and skills, so external tools can be added without code.
- **Scan plan:** set a repository to `Manual` to override depth and file caps.

## Running it

`compose.yaml` runs one container on port 18081 with SQLite at `/data/opendeepwiki.db` (`compose.pgsql.yaml` for PostgreSQL). Models are set with `WIKI_CATALOG_*`, `WIKI_CONTENT_*` and optional `WIKI_TRANSLATION_*` variables. Local imports must sit under `RepositoryAnalyzer__AllowedLocalPathRoots`.

This site built the backend from the pinned source and ran it headless on every repository in our catalog. Notes from that run:

- On startup, the seeder marks a preset active only if it has `defaultEnabled`, and the built-in `openrouter` preset has none ([AiProviderPresetSeeder.cs#L183](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/AI/AiProviderPresetSeeder.cs#L183)). The `WIKI_*` migration then matches the built-in row, stores the key and does not activate it ([AiConfigurationMigrationService.cs#L161-L167](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/AI/AiConfigurationMigrationService.cs#L161-L167)). Every generation failed with "AI provider is not available" until we logged in as the seeded admin and enabled the provider with `PUT /api/admin/tools/ai-providers/{id}`.
- A seeded admin account with a hard-coded password is created on first start ([DbInitializer.cs#L129-L130](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Infrastructure/DbInitializer.cs#L129-L130)), and the compose file has a default JWT secret. Change both before you expose the port.
- The catalog agent saw at most about 500 files of tree per repo. The mind map is a separate job (`MindMapWorker`) and can take a large share of tokens: 508K input tokens for rakazo, against 92K for its catalog.
- Pages come from per-page HTTP calls (`/api/v1/repos/{owner}/{repo}/tree`, then `/docs/{slug}`). The `/export` zip is rate-limited.

Measured with DeepSeek V4 Flash on the same three repos as [deepwiki-open](/p/deepwiki-open/) (wall time including about 25 s of queue wait; cost is server token counts at list price):

| Repo | OpenDeepWiki | deepwiki-open |
|---|---|---|
| llm-scraper | 9 pages, 300 s, $0.11 | 11 pages, 285 s |
| OpenBot | 61 pages, 1372 s, $0.70 | 12 pages, 798 s |
| rakazo | 17 pages, 401 s, $0.21 | 10 pages, 512 s |

More than 90% of OpenDeepWiki's input tokens were prompt-cache hits, which keeps the agent loop cheap. Page count follows repository size, so cost scales with the repo, not with a fixed outline.

## Strengths and caveats

- **Strength: the model reads code itself.** Pages are built from files the agent chose and read, and `SourceFiles` records exactly which ones.
- **Strength: incremental updates.** A new commit triggers `IncrementalUpdateAsync()` on changed files instead of a full rebuild.
- **Strength: operations.** Leases, heartbeats, stale-work recovery, per-page timeouts, resume of half-finished runs, and per-operation token accounting.
- **Strength: wide output.** Multi-language docs, mind map, chat, and MCP per repository.
- **Caveat: tight budgets.** With 6 source calls and 240-line reads per page, deep pages on large modules rely on few excerpts. Raise the budgets if pages are thin.
- **Caveat: weak per-repo doc search.** The repo-scoped MCP `SearchDoc` is a literal substring match, so a full-sentence question often finds nothing. The global endpoint's token-scored search does better.
- **Caveat: heavy to run.** It needs a .NET service, a database, an admin login and provider setup in the DB. Several bundled features (chat-platform bots, Graphify, recommendations) are outside the core use case.
- **Caveat: failed pages stay missing.** A page that fails or hits the 60-minute timeout is logged and skipped. The run still completes, and the gap is filled only by a later regenerate (`ThrowOnPartialFailure` is off by default).
- **Caveat: legacy leftovers.** `Directory.Packages.props` still pins Semantic Kernel, MySQL and SQL Server packages, and `CLAUDE.md` mentions legacy KoalaWiki projects. The service itself references Microsoft.Agents.AI, targets `net10.0`, and ships migrations only for SQLite and PostgreSQL. The deepwiki-open wiki we generated for this repo described the old KoalaWiki design (Semantic Kernel, a five-step "KoalaWarehouse" pipeline, four databases), so check generated docs against the code.

*Sources: code at d33113c, deepwiki-open wiki (12 pages), verified Q&A.*

## How AIDotNet/OpenDeepWiki answers the Open-source DeepWiki questions

### How is a repository ingested and chunked? (answered)

**Repository sources and hosts.** Repositories are imported via `RepositorySource.Parse()` in `RepositoryAnalyzer` (src/OpenDeepWiki/Services/Repositories/RepositoryAnalyzer.cs:82) which supports three `RepositorySourceType` values: `Git` (remote Git URL), `Archive` (uploaded ZIP), and `LocalDirectory` (approved local paths). Git repos use LibGit2Sharp for clone/pull, with retries (default 3 attempts, 1 s delay). Local directories are either symlinked or copied under `REPOSITORIES_DIRECTORY`. ZIP archives are extracted via `ZipFile.ExtractToDirectory`. The config object `RepositoryAnalyzerOptions` (line 15) specifies the base clone directory (`/data` by default) and an allowlist `AllowedLocalPathRoots` for local imports.

**File filters and scan plans.** The `RepositoryFileFilter` class (src/OpenDeepWiki/Agents/Tools/RepositoryFileFilter.cs:6) parses `.gitignore` files recursively and applies their patterns via `GitIgnorePatternToRegex` (line 181). Hidden paths (`.`-prefixed directories) are automatically skipped, and binary files are filtered by extension. The scan plan resolver `RepositoryScanPlanResolver` (src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs:12) profiles the working directory — counting files/directories by depth, extension distributions, key files (READMEs, Makefiles, etc.) — and then applies rule-based heuristics (line 87–125). For repos with >20000 total files, tree depth is clamped to 2, max total files to 350; for 5000–20000, depth goes to 3; under 5000, depth 4. Max tree nodes budget is 1500, max files per directory 30, max total files 800.

**Large-repo limits.** The hardcoded budgets act as safety valves: `MaxTreeNodeBudget = 1500`, `MaxFilesPerDirectoryBudget = 30`, `MaxTotalFileBudget = 800` (line 15–19). These are used in `Clamp()` calculations in `DecideByRules()` (line 87). The plan can also be set to `Manual` mode where overrides from the repository entity are used, or `Auto` where it is computed and persisted.

**Chunking strategy.** This project does NOT chunk source code into fixed-size blocks for embedding. Instead the entire workspace (or a bounded directory/file listing) is provided to an AI agent via tool calls. The file listing depth, max nodes, and max files per directory are the chunking controls — they limit how much of the tree is presented to the LLM in a single catalog/content generation call.

> **Editor's note.** Correction: the effective file-tree cap is `min(MaxTotalTreeFiles, 500)`, dropping to 350 above 20,000 files. The 800 and 1,500 values are only outer clamps. During generation the tool budget caps each read at 240 lines, each listing at 20 entries and each grep at 12 results.

Citations: [src/OpenDeepWiki/Services/Repositories/RepositoryAnalyzer.cs:82-240](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryAnalyzer.cs#L82-L240) · [src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs:12-125](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs#L12-L125) · [src/OpenDeepWiki/Agents/Tools/RepositoryFileFilter.cs:6-258](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/Tools/RepositoryFileFilter.cs#L6-L258) · [src/OpenDeepWiki/Services/Repositories/RepositoryScanPlan.cs:1-48](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryScanPlan.cs#L1-L48)

### How is retrieval (RAG) implemented? (answered)

There is no traditional RAG pipeline — no vector embeddings, no vector database, and no top-k similarity search. Instead, OpenDeepWiki uses **agentic file and document reading** via AI tool calls, implemented in two key classes.

**Pre-generated document retrieval.** The `DocReadTool` (src/OpenDeepWiki/Agents/Tools/DocReadTool.cs:14) exposes `ReadDocumentAsync(path)` and `ListDocumentsAsync()` to the AI chat agent. It looks up the repository, branch, and language in EF Core (DocCatalog + DocFile entities) and returns cached Markdown content. Each document record stores a `SourceFiles` JSON array (the source code files that were analyzed to generate it). The `ChatDocReaderTool` (src/OpenDeepWiki/Agents/Tools/ChatDocReaderTool.cs:15) provides a more granular `ReadAsync(path, startLine, endLine)` capped at 200 lines per request, with dynamic tool descriptions that embed the available-document catalog.

**Source-code retrieval.** The `GitTool` (src/OpenDeepWiki/Agents/Tools/GitTool.cs:13) exposes three AI tools: `ReadFile(path, offset, limit)` (line 180), `ListFiles(glob, maxResults)` (line 493), and `Grep(pattern, glob, ...)` (line 283). These let the agent read arbitrary source files (respecting `.gitignore` and hidden-path filtering), glob for project structure, and regex-search across files. Lines are capped at 2000 characters; reads default to 2000 lines; grep results default to 50 with 2 context lines. Binary files are skipped by extension.

**How it reaches the prompt.** In `ChatAssistantService.StreamChatAsync()` (src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:372), tools are constructed at request time: a `GitTool` is initialized from the cloned repository path on disk, a `ChatDocReaderTool` loads the document catalog from the database, and optionally MCP/skill tools. The system prompt (line 963–1138) instructs the agent to first `ReadDoc` for documentation, check the `sourceFiles` field, then use `ReadFile`/`Grep` for deeper source code investigation. The `DocumentSourceToolBudget` (src/OpenDeepWiki/Agents/Tools/DocumentSourceToolBudget.cs) caps how many source-discovery tools (ListFiles/Grep/ReadFile) the agent can use per page.


Citations: [src/OpenDeepWiki/Agents/Tools/DocReadTool.cs:14-260](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/Tools/DocReadTool.cs#L14-L260) · [src/OpenDeepWiki/Agents/Tools/ChatDocReaderTool.cs:15-292](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/Tools/ChatDocReaderTool.cs#L15-L292) · [src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:372-1138](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs#L372-L1138)

### How is the wiki structure (table of contents) determined? (answered)

The wiki structure (table of contents) is generated by an **LLM agent using the catalog-generator prompt** (src/OpenDeepWiki/prompts/catalog-generator.md). `WikiGenerator.GenerateCatalogAsync()` (src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs:245) is the entry point. It first calls `CollectRepositoryContextAsync()` to gather the project's directory tree (TOON format), README content, key files, and entry points. Then it loads the `catalog-generator` prompt template, creates a `GitTool` (for source exploration), a `CatalogTool` (wrapping `CatalogStorage` to write the catalog), and a `DocumentSourceToolBudget` that limits source tool calls to the configured `MaxCatalogSourceToolCalls` (default 4). The AI agent is instructed to design a **DeepWiki-style information architecture** — top-level topic domains with independently readable deep-dive leaf pages — not a file/class listing.

**Inputs.** The agent receives a runtime context containing the directory tree, README content, key files, entry points, and project type. It uses the tools sparingly for confirmation reads/greps before writing the catalog.

**Output format.** The catalog is JSON via the `WriteCatalog` tool (`CatalogTool.WriteAsync`), following:
```json
{"items": [{"title": "...", "path": "lowercase-hyphen-path", "order": N, "children": []}]}
```
Each leaf node gets a `DocFileId` pointing to generated content. Parent nodes are navigation groupings only.

**Design rules.** The prompt (line 25–75) emphasizes right-sized coverage (no fixed counts), no over-compression (many independent systems must not be hidden inside few chapters), verifying via source tools first, and never fabricating entries. Page titles should be in the runtime target language. The `CatalogStorage` class (src/OpenDeepWiki/Services/Wiki/CatalogStorage.cs:12) persists the catalog to `DocCatalog` entities in the database, keyed by `BranchLanguageId`. Parent-child relationships are tracked via `ParentId` and an `Order` field.


Citations: [src/OpenDeepWiki/prompts/catalog-generator.md:1-215](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/prompts/catalog-generator.md#L1-L215) · [src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs:245-362](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs#L245-L362) · [src/OpenDeepWiki/Services/Wiki/CatalogStorage.cs:1-60](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/CatalogStorage.cs#L1-L60)

### How are individual pages generated? (answered)

**Per-page prompts.** Each wiki page is generated by an LLM agent using the `content-generator` prompt (src/OpenDeepWiki/prompts/content-generator.md). `WikiGenerator.GenerateDocumentsAsync()` (src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs:366) iterates over all leaf catalog items. For each item, `GenerateDocumentContentAsync()` (line 796) builds a user message with the catalog path/title, a git-based File Reference Base URL, and budget limits. The agent receives a `GitTool` (capped at `MaxDocumentSourceToolCalls`, default 6), a `DocTool` (for WriteDoc/AppendDoc/EditDoc), and a tool budget (`MaxDocumentToolCalls`, default 14; `MaxDocumentAppendOperations`, default 4).

**Source-file citations.** The agent is instructed to read 2-4 key source files, extract code snippets, and embed them with Markdown blockquote source links using the runtime File Reference Base URL (line 843–848). The `DocTool` automatically tracks which files were read; these are stored as a JSON array in `DocFile.SourceFiles` on the entity.

**Mermaid diagrams.** The prompt mandates at least one Mermaid diagram per page (architecture/flowchart/sequence/class/ER). Extensive Mermaid syntax rules are embedded in the prompt (line 850–900), including node ID naming, subgraph `sg_` prefix requirements, and edge label quoting — reflecting real renderer constraints encountered in production.

**Parallelism.** `GenerateDocumentsAsync()` uses `Parallel.ForEachAsync` with `MaxDegreeOfParallelism` from `WikiGeneratorOptions.ParallelCount` (default 5, line 60). Before parallel execution, it "warms" the prompt cache by generating the first document sequentially (line 544–559), then processes the remaining items in parallel. Timed-out or failed documents are logged but don't stop the batch (line 497–542) unless `ThrowOnPartialFailure` is set. Each document has a timeout of `DocumentGenerationTimeoutMinutes` (default 60).

**Caching and regeneration.** Documents that already have persisted content via `DocFileId` are skipped (line 414–421). `RegenerateDocumentAsync()` (line 611) allows forced regeneration of a single path. Incremental updates (via `IncrementalUpdateAsync`, line 660) use the `incremental-updater` prompt with a restricted toolset (no `WriteCatalog`) to edit only affected documents based on `git diff` results between commits.


Citations: [src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs:366-1100](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs#L366-L1100) · [src/OpenDeepWiki/prompts/content-generator.md:1-1217](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/prompts/content-generator.md#L1-L1217) · [src/OpenDeepWiki/Services/Wiki/WikiGeneratorOptions.cs:1-241](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGeneratorOptions.cs#L1-L241)

### How are model providers configured? (answered)

**Provider model.** Providers are configured through database entities `AiProviderConfig` and `AiModelConfig`, or through model configs (`ModelConfig`). The `AiProviderResolver` (src/OpenDeepWiki/Services/AI/AiProviderResolver.cs:52) resolves a provider+model pair into a `ResolvedAiModel` with endpoint, API key, request type, token pricing, and capabilities. It queries the database for active non-deleted configurations.

**Supported provider types.** `AiProviderResolver.ParseRequestType()` (line 199–208) maps normalized provider type strings to `AiRequestType` enum values: `OpenAI`, `AzureOpenAI`, `OpenAIResponses`, `Anthropic`, `DeepSeekOpenAI`. The `NormalizeProviderType()` function (line 137–153) handles aliases: `azure`, `azure-openai` → `AzureOpenAI`; `anthropic`, `claude` → `Anthropic`; `deepseek`, `deepseek-openai-chat` → `DeepSeekOpenAI`; `openai-chat`, `openai` → `OpenAI`. The `AgentFactory` (src/OpenDeepWiki/Agents/AgentFactory.cs:60) constructs the appropriate HTTP client per type — for Anthropic it uses an Anthropic SDK client, for DeepSeek a `DeepSeekOpenAIChatClient`, for OpenAI the OpenAI SDK (chat or responses API).

**Per-stage model selection.** The `WikiGeneratorOptions` (src/OpenDeepWiki/Services/Wiki/WikiGeneratorOptions.cs:9) has **separate fields** for each stage: `CatalogModel` (for catalog and mind map generation), `ContentModel` (for document content), `TranslationModel` (optional fallback to ContentModel). Each model field can be paired with a matching `*ProviderId` (e.g., `CatalogProviderId`, `ContentProviderId`) to bind a specific provider. `WikiGeneratorOptionsConfigurator` (src/OpenDeepWiki/Services/Wiki/WikiGeneratorOptionsConfigurator.cs:5) wires environment variables: `WIKI_LANGUAGES`, `WIKI_MAX_CONCURRENT_GENERATIONS`, plus per-stage config via the `WikiGenerator` config section.

**OpenAI-compatible endpoints.** Any provider can use an arbitrary `BaseUrl` (line 210–226). If the base URL is empty, defaults are assigned per type (Anthropic → `https://api.anthropic.com`, DeepSeek → `https://api.deepseek.com/v1`, others → `https://api.openai.com/v1`). This makes any OpenAI-compatible endpoint (vLLM, Ollama, etc.) usable by setting the request type to `OpenAI`.

**Built-in presets.** `AiProviderPresetCatalog` (src/OpenDeepWiki/Services/AI/AiProviderPresetCatalog.cs:12) loads embedded JSON presets from `Services.AI.BuiltinProviderPresets.json` containing provider definitions like name, type, default base URL, default model list, and pricing. The presets support seeding on first run via `AiProviderPresetSeeder`.

**Overrides.** Each provider and model carries `RequestOverridesJson` and `ThinkingConfigJson`, allowing per-invocation passthrough of provider-specific parameters. The `AiRequestOptions` (src/OpenDeepWiki/Agents/AgentFactory.cs:26) exposes `SupportsThinking`, `ThinkingEnabled`, and all override fields, which are applied in the handler chain.


Citations: [src/OpenDeepWiki/Services/AI/AiProviderResolver.cs:1-228](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/AI/AiProviderResolver.cs#L1-L228) · [src/OpenDeepWiki/Agents/AgentFactory.cs:17-80](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/AgentFactory.cs#L17-L80) · [src/OpenDeepWiki/Services/Wiki/WikiGeneratorOptions.cs:1-241](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Wiki/WikiGeneratorOptions.cs#L1-L241)

### How is interactive Q&A / chat implemented? (answered)

**Chat over the repo.** The chat assistant is served through `ChatAssistantEndpoints` (src/OpenDeepWiki/Endpoints/ChatAssistantEndpoints.cs:14). The primary endpoint is `POST /api/v1/chat/stream` (line 150), which accepts a `ChatRequest` containing message history, model ID, and a `DocContextDto` (owner, repo, branch, language, current document path, catalog menu). The response is Server-Sent Events (SSE) with `text/event-stream` content type and `X-Accel-Buffering: no` to disable nginx buffering.

**Streaming implementation.** `ChatAssistantService.StreamChatAsync()` (src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:372) creates an agent session via `Microsoft.Agents.AI`. It uses `agent.RunStreamingAsync()` to yield token-by-token SSE events of types: `content` (text chunks), `thinking` (reasoning tokens from models that support it, e.g. DeepSeek R1), `tool_call` (with incremental or complete arguments), `tool_result`, `done` (final token counts), and `error`. The tool-call handling supports both **OpenAI streaming format** (`StreamingChatCompletionUpdate` with `ToolCallUpdates`, line 590–664) and **Anthropic streaming format** (`RawMessageStreamEvent` with `content_block_start/delta/stop` events, line 666–803), including `thinking_delta` and `input_json_delta` for extended thinking models.

**Agent tools.** The chat agent receives `ReadDoc` (line 503) for reading generated wiki pages with line ranges, `ReadFile`/`ListFiles`/`Grep` from `GitTool` (line 479) for source code access when the repository clone is available, plus configurable MCP tools and Skill tools. The system prompt (line 963–1138) instructs a structured workflow: Understand → Gather (using tools) → Analyze → Respond, with source-code verification required when docs are unclear.

**Conversation memory.** Session state is tracked in the database via `ChatSessionImpl` and `SessionManager` (src/OpenDeepWiki/Chat/Sessions/). Messages are persisted through a `DatabaseMessageQueue` with dead-letter handling. Usage accounting records token consumption via `AiUsageAccounting` per session.

**Deep-research mode.** The system prompt includes `internal_thinking` instructions (line 998–1022) that encourage the agent to reason step-by-step before responding, using  tags for structured deliberation.

**MCP exposure.** The chat is also available through MCP endpoints (`/api/mcp` and `/api/mcp/{owner}/{repo}`, per README), where the same agent tools are exposed via the MCP protocol. Repository-scoped MCP endpoints bind the tool context to a specific repo. The `McpToolConverter` (src/OpenDeepWiki/Services/Chat/McpToolConverter.cs) converts MCP provider configs into AI tools at request time.


Citations: [src/OpenDeepWiki/Endpoints/ChatAssistantEndpoints.cs:1-245](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Endpoints/ChatAssistantEndpoints.cs#L1-L245) · [src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:186-300](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs#L186-L300) · [src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:463-827](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs#L463-L827) · [src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs:963-1138](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Chat/ChatAssistantService.cs#L963-L1138)
