# AsyncFuncAI/deepwiki-open

> FastAPI service and Next.js UI that embed a repo into a FAISS index, then generate a cited, Mermaid-heavy wiki with one prompt per page.

- Category: [Open-source DeepWiki](https://llms-technical-reviews.com/open-source-deepwiki/)
- Repository: https://github.com/AsyncFuncAI/deepwiki-open (reviewed at commit `d92819a9c9f3b99416e3580ff235fc9d3adf8b89`, 2026-09-03)
- Stars: 18128 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/deepwiki-open/

## Overview

deepwiki-open is a self-hosted wiki generator built as a classic RAG pipeline. A Python FastAPI backend (`api/`) clones or reads a repository, splits every eligible file into word chunks, embeds them, and saves a FAISS-ready index as a pickle. It then asks an LLM for a wiki outline in XML and writes each page with a separate LLM call. A Next.js frontend (`src/`) renders the wiki and offers chat over the same index.

The LLM never reads files on its own. Everything it sees comes from the prompt: the file tree and README for the outline, then the top-20 retrieved chunks for each page or chat turn. This keeps the design simple and cheap to run. It also means page quality depends on what the retriever happens to return, not on what the model would have chosen to read.

It fits a single team that wants a local DeepWiki for GitHub, GitLab, Bitbucket or local repositories, with a free choice of model provider (Google, OpenAI, OpenRouter, Ollama, Bedrock, Azure, DashScope, LiteLLM, Anthropic). It has no user accounts, no database and no incremental update. The output is one JSON cache file per repository and language.

## Architecture

```mermaid
flowchart TD
    UI["Next.js UI (src/)"] -->|"POST /wiki/tasks"| WR["routers/wiki.py"]
    UI -->|"WS /ws/chat"| CR["routers/chat.py"]
    WR --> TR["TaskRegistry (tasks.py)"]
    TR --> GEN["generate_repo_wiki()"]
    GEN --> REPO["Repo: clone or local path"]
    GEN --> IDX["prepare_repo_index()"]
    IDX --> PIPE["split + embed (rag/pipeline.py)"]
    PIPE --> PKL[("LocalDB .pkl")]
    GEN --> STRUCT["_determine_structure()"]
    GEN --> PAGES["_generate_pages()"]
    STRUCT --> RC["research_chat()"]
    PAGES --> RC
    CR --> RC
    RC --> FAISS["FAISSRetriever top_k=20"]
    FAISS --> PKL
    RC --> CS["ChatStreamer.create(provider)"]
    CS --> LLM["LLM provider"]
    PAGES --> CACHE[("wikicache JSON")]
```

| Component | Path | Role |
|---|---|---|
| App entry | [api/main.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/main.py#L64-L72) | Mounts the system, auth, repo, wiki, chat and codemap routers |
| Wiki task API | [api/routers/wiki.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/routers/wiki.py#L218-L302) | Submit, list, poll and SSE-stream generation tasks |
| Task runner | [api/services/wiki/tasks.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L140-L209) | In-memory registry, semaphore, state machine |
| Repository access | [api/repository.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/repository.py#L140-L229) | Shallow clone with token injection, or use a local path in place |
| File filter | [api/config.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config.py#L467-L516) | `iterate_files()`: extension allow-list plus include/exclude rules |
| Indexing | [api/rag/pipeline.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/pipeline.py#L174-L275) | Line-tracking splitter, embedder, LocalDB pickle |
| Retrieval + prompt | [api/services/research.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/research.py#L60-L311) | `research_chat()`: one path for outline, pages and chat |
| Providers | [api/chat/_stream.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/chat/_stream.py#L29-L48) | `ChatStreamer` registry, one subclass per provider |
| Prompts | [api/services/wiki/prompts.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/prompts.py#L213-L260) | Outline (XML) and page prompts |
| Post-processing | [api/services/wiki/content.py](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/content.py#L84-L151) | Turns `[path:10-25]()` citations into real host links |

## How a request flows

1. The client calls `POST /wiki/tasks`. [`submit_wiki_task`](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/routers/wiki.py#L218-L228) hands a `WikiTask` to `TaskRegistry.submit()`. If a task for the same repo is active, the caller joins it. If a cache file already exists for that repo and language, the call returns `from_cache` and nothing is regenerated ([tasks.py#L161-L190](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L161-L190)).
2. `TaskRegistry._run()` waits on a semaphore sized by `DEEPWIKI_MAX_CONCURRENT_WIKI_TASKS`, then runs [`generate_repo_wiki()`](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L212-L239).
3. **Indexing.** If no `.pkl` exists, `prepare_repo_index()` builds a `RAG` object and calls `aprepare_retriever()`. `Repo.download()` does a `--depth=1 --single-branch` clone; local paths are used as they are ([repository.py#L191-L222](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/repository.py#L191-L222)). Files come from `iterate_files()`. Files above `MAX_EMBEDDING_TOKENS * 10` tokens are skipped ([pipeline.py#L112-L130](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/pipeline.py#L112-L130)). The rest are split, embedded and saved with `LocalDB.save_state()`.
4. **Outline.** [`_determine_structure()`](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L315-L358) reads the file list and the shortest-path README, and builds the outline prompt with `build_structure_prompt()`. It sends the prompt through `research_chat()` and parses the reply with `parse_wiki_structure()`. Any failure here fails the whole task.
5. **Pages.** [`_generate_pages()`](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L288-L312) runs `_generate_page()` for every page under a second semaphore (`DEEPWIKI_WIKI_PAGE_CONCURRENCY`, default 1). Each page is retried `WIKI_PAGE_RETRIES` times, then replaced by an error placeholder, so one bad page does not fail the wiki.
6. Inside [`research_chat()`](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/research.py#L144-L199), the last user message is the retrieval query. The top-20 chunks are grouped by file and annotated with `[lines a-b]`. `prompt_builder()` wraps them in `<START_OF_CONTEXT>` tags ([_prompts.py#L6-L31](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/chat/_prompts.py#L6-L31)).
7. `ChatStreamer.create()` picks the provider class and streams the answer. `post_process_wiki_content()` rebuilds the `<details>` source list and resolves citations. `save_wiki_cache()` writes the JSON file. Progress is visible through `GET /wiki/tasks/{id}/stream` (SSE, one event per second).

## Key components

### Retrieval index

The splitter config is `split_by: word`, `chunk_size: 350`, `chunk_overlap: 100`, and the retriever uses `top_k: 20` (`api/config/embedder.json`). `LineTrackingTextSplitter` finds each chunk back in its parent text and stores 1-based `start_line`/`end_line`, so the model can cite real line numbers ([pipeline.py#L174-L208](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/pipeline.py#L174-L208)). The index is a pickle at `<root>/databases/<repo>.pkl`, and its existence is the only "already indexed" check ([pipeline.py#L163-L171](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/pipeline.py#L163-L171)). A `FAISSRetriever` is built in memory from the stored vectors ([rag.py#L288-L296](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/rag.py#L288-L296)).

### One prompt path for everything

The outline, every page, and every chat turn all go through `research_chat()`. Two details matter. First, the retrieval query for a page is the whole page prompt, which is mostly fixed instructions plus the title and file links. Second, if that last message is over `MAX_INPUT_TOKENS = 7500` tokens, retrieval is skipped ([research.py#L22-L23](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/research.py#L22-L23), [#L64-L73](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/research.py#L64-L73)). For most real repos the outline prompt carries the full file list and README, so it crosses that limit and the outline is built from the tree and README only.

### Page prompt and citations

[`build_page_prompt()`](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/prompts.py#L25-L40) tells the model it has "the full content" of the listed files and must cite at least five of them. In fact it only receives retrieved chunks. Citations are written as `Sources: [path:10-25]()` with empty parentheses ([prompts.py#L108-L118](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/prompts.py#L108-L118)). `post_process_wiki_content()` resolves them against the page's file list, then generic paths, then bare file names, into GitHub, GitLab or Bitbucket URLs.

### Providers

Each `ChatStreamer` subclass registers itself by its `provider` attribute ([_stream.py#L29-L48](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/chat/_stream.py#L29-L48)). `get_model_config()` reads `generator.json` and falls back to the provider's default model's parameters for unknown model names ([config.py#L364-L417](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config.py#L364-L417)). The OpenRouter client always sends `stream: False` and sets a 60 s timeout, even though the streamer asks for streaming ([openrouter.py#L140-L157](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/clients/openrouter.py#L140-L157)).

### Chat, Deep Research and Codemap

`/ws/chat` ([chat.py#L20-L66](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/routers/chat.py#L20-L66)) streams `research_chat()` output as text frames. History is sent by the client on each request and replayed as `<turn>` blocks. Deep Research is a prompt switch on `research_iteration` (plan, updates, final answer at iteration 5); the frontend drives the loop. Codemap (`api/services/codemap.py`) is a separate two-call pipeline that returns NDJSON.

## Extending it

- **New provider:** subclass `ChatStreamer` with a `provider` name and add a block in `api/config/generator.json`. Registration is automatic.
- **Embedder:** `DEEPWIKI_EMBEDDER_TYPE` (openai, google, ollama, bedrock) selects a block in `embedder.json`. Any OpenAI-compatible embedding endpoint works through `OPENAI_BASE_URL`.
- **Filters:** `api/config/repo.json` holds the extension lists and default excluded dirs. Requests can add `excluded_dirs`, `excluded_files`, `included_dirs`, `included_files`.
- **Prompts:** outline and page prompts are plain f-strings in `api/services/wiki/prompts.py`; chat and research prompts are in `api/prompts.py`.

## Running it

The repo ships a [docker-compose.yml](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/docker-compose.yml) that runs the API on 8001 and the UI on 3000 and mounts `~/.adalflow` for clones and indexes. You need one generator key (Google by default) and one embedder key; the default embedder is OpenAI `text-embedding-3-small`.

This site ran the API headless on every repository in our catalog. Notes from that run:

- We drove it with `POST /wiki/tasks` and a local repo path, read-only mounted, so it never cloned from GitHub.
- Without an OpenAI key you can point the OpenAI embedder at any OpenAI-compatible endpoint through `OPENAI_BASE_URL` (we used `openai/text-embedding-3-small` at 256 dims) and a replaced `embedder.json`.
- The OpenRouter client's 60 s non-streaming timeout killed whole-repo outline calls. A failed call comes back as error text, the XML parse fails, and the whole task fails. We patched the timeout to 900 s.
- `MAX_CONCURRENT_WIKI_TASKS` defaults to `os.cpu_count() // 2` ([tasks.py#L61-L66](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L61-L66)). Python reports host cores, not the container CPU limit, so on a large host it would launch far more tasks than a small container can run; set it explicitly.
- Tasks live only in memory. A restart loses running work, and finished tasks drop out of the registry after 300 s; after that, only the cache file shows the result.

Measured on the same three repos, with DeepSeek V4 Flash for both tools and 3 pages in parallel here: llm-scraper 11 pages in 285 s, OpenBot 12 pages in 798 s, rakazo 10 pages in 512 s. deepwiki-open records no token usage, so we have no per-wiki cost for it. The [OpenDeepWiki](/p/opendeepwiki/) page has its numbers.

## Strengths and caveats

- **Strength: simple and portable.** One Python process, files on disk, no database. Nine provider classes and Ollama make fully local runs possible.
- **Strength: good citation plumbing.** Line-tracked chunks plus `post_process_wiki_content()` give pages real `#Lx-Ly` links.
- **Strength: fixed, predictable cost.** One outline call plus one call per page, with a bounded prompt.
- **Caveat: retrieval-bound pages.** The model sees only 20 chunks per page, retrieved with the page prompt as the query, even though the prompt claims full file access.
- **Caveat: outline skips RAG on large repos** because of the 7500-token check on the query, and the file list is sent as a Python list literal.
- **Caveat: argument-order bug.** `_determine_structure()` passes `(excluded_dirs, excluded_files, included_dirs, included_files)` by position, but [`read_repo_file_tree()`](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/structure.py#L20-L48) expects `(included_files, included_dirs, excluded_files, excluded_dirs)` ([tasks.py#L328-L335](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L328-L335)). A request with exclusions therefore builds the outline in inclusion mode. The index itself uses keyword arguments and is not affected.
- **Caveat: no incremental update.** A cached wiki is returned as-is; to refresh it you delete the cache and regenerate everything.
- **Caveat: no output check.** `stream_and_fallback()` yields provider errors as text, so a failed page can be saved as an error string instead of being retried. Nothing checks the page shape either: when we ran deepwiki-open on its own repo, the "Project Overview" page came back as about 600 lines of the model's reasoning, and it was saved as a finished page.
- **Caveat: leftover frontend paths.** The UI still contains a `POST /api/wiki_cache` call for saving a finished wiki, but the backend only defines `GET` and `DELETE` for that route ([wiki.py#L139-L165](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/routers/wiki.py#L139-L165)). The backend task writes the cache itself.

*Sources: code at d92819a, deepwiki-open wiki (12 pages), verified Q&A.*

## How AsyncFuncAI/deepwiki-open answers the Open-source DeepWiki questions

### How is a repository ingested and chunked? (answered)

A repository is ingested via the `Repo` class in `api/repository.py` (line 140). It supports GitHub, GitLab, Bitbucket, and local paths — each with its own `_clone_from_*` function that handles access tokens via URL injection (`oauth2:token@` for GitLab, `token@` for GitHub, `x-bitbucket-api-token-auth` for Bitbucket). Clones use `git.Repo.clone_from` with `--depth=1 --single-branch` for shallow clones. Local paths skip cloning entirely and are used in-place. The cloned repo path is `~/.deepwiki/repo/`.

File filtering is centralized in `api/config.py`'s `iterate_files()` (line 467). It walks the repo directory and applies: (1) extension whitelisting from `code_extensions` and `doc_extensions` config arrays, (2) exclusion mode using default `file_filters.excluded_dirs` (like `node_modules`, `.venv`, `.git`, `dist`) merged with request-level exclusions, or (3) inclusion mode when `included_dirs`/`included_files` are specified. Default excluded dirs are extensive (line 3-40 of `repo.json`).

Chunking happens in the RAG pipeline, not at ingestion. `api/rag/pipeline.py`'s `LineTrackingTextSplitter` (line 174) uses adalflow's `TextSplitter` with `split_by: word`, `chunk_size: 350`, `chunk_overlap: 100` (from `embedder.json` line 36-41). Each chunk gets annotated with its 1-based `start_line`/`end_line` within the original file. Files exceeding `MAX_EMBEDDING_TOKENS * 10` (~82K tokens) are skipped (line 126). The pipeline then embeds each chunk and saves to a local `LocalDB` pickle file at `~/.adalflow/databases/` (line 247-274). Large-repo limits are configurable via `MAX_CONCURRENT_WIKI_TASKS` (defaults to half CPU cores) and `WIKI_PAGE_CONCURRENCY` (default 1) in `api/services/wiki/tasks.py` (lines 62-68).


Citations: [api/config.py:467-516](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config.py#L467-L516) · [api/config/repo.json:3-150](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config/repo.json#L3-L150) · [api/rag/pipeline.py:174-275](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/pipeline.py#L174-L275) · [api/services/wiki/tasks.py:62-68](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L62-L68) · [api/config/embedder.json:33-41](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config/embedder.json#L33-L41)

### How is retrieval (RAG) implemented? (answered)

Retrieval is implemented via a FAISS-based RAG pipeline using the `adalflow` framework. The core class is `RAG` in `api/rag/rag.py` (line 134), which wraps a `FAISSRetriever` (line 290) configured with `top_k: 20` (from `embedder.json`). Documents are embedded using configurable embedders: OpenAI `text-embedding-3-small` (256 dims) by default, Google `gemini-embedding-001`, Ollama `nomic-embed-text`, or Bedrock `amazon.titan-embed-text-v2:0` — selected via the `DEEPWIKI_EMBEDDER_TYPE` env var (config.py line 65).

The retrieval pipeline in `research_chat()` (`api/services/research.py`, line 60) first prepares the retriever index (loading from a `.pkl` pickle file or creating it fresh), then calls `rag.acall(query)` which invokes the FAISS retriever. Retrieved documents are grouped by file path and each chunk is annotated with its line range (line 178-184). The grouped context is assembled into a structured prompt that includes file-path headers and line-numbered chunks inside `<START_OF_CONTEXT>` / `<END_OF_CONTEXT>` tags (line 146-192).

The context is then injected into `prompt_builder()` (`api/chat/_prompts.py`, line 6), which builds a final prompt combining the system prompt, conversation history, context block, and user query. The response streamer (`ChatStreamer` in `_stream.py`) sends the full prompt to the LLM and streams back chunks. If the input exceeds ~7500 tokens, RAG context is omitted and a fallback note is inserted instead (line 27-29 of `_prompts.py`). The `Memory` class in `rag.py` (line 79) accumulates conversation turns for multi-turn dialogue.


Citations: [api/rag/pipeline.py:170-208](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/pipeline.py#L170-L208) · [api/services/research.py:60-311](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/research.py#L60-L311) · [api/chat/_prompts.py:6-31](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/chat/_prompts.py#L6-L31) · [api/config/embedder.json:1-41](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config/embedder.json#L1-L41) · [api/config.py:65-65](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config.py#L65-L65)

### How is the wiki structure (table of contents) determined? (answered)

The wiki structure (table of contents) is determined by an LLM call. In `api/services/wiki/tasks.py`, `_determine_structure()` (line 315) orchestrates the process. First it reads the cloned repo's file tree and README via `read_repo_file_tree()` in `structure.py` (line 20), which returns sorted file paths filtered through `iterate_files()` plus the README content. The repo's default branch is detected with `git rev-parse --abbrev-ref HEAD` (line 51).

These inputs are fed into `build_structure_prompt()` in `prompts.py` (line 213). The prompt constructs an XML schema asking the LLM to output a `<wiki_structure>` block containing a title, description, and individual `<page>` elements — each with an id, title, importance (high/medium/low), related pages, and a list of `<file_path>` citations. For "comprehensive" wikis, it also includes `<sections>` grouping pages and supporting nested subsections (line 135-185); "concise" omits sections and targets 4-6 pages instead of 8-12 (line 187-209).

The LLM's raw XML response is parsed by `parse_wiki_structure()` in `structure.py` (line 179). This function is robust: it strips markdown fences, escapes bare `&` characters, handles truncated responses (missing `</wiki_structure>`) by salvaging complete blocks, and falls back to regex-based extraction if strict `xml.etree.ElementTree` parsing fails. The result is a `WikiStructureModel` containing `WikiPage` and `WikiSection` objects (schemas in `api/schemas/wiki.py`).


Citations: [api/services/wiki/tasks.py:315-358](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L315-L358) · [api/services/wiki/prompts.py:187-260](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/prompts.py#L187-L260) · [api/services/wiki/structure.py:1-250](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/structure.py#L1-L250) · [api/schemas/wiki.py:1-43](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/schemas/wiki.py#L1-L43)

### How are individual pages generated? (answered)

Individual wiki pages are generated by `_generate_pages()` in `api/services/wiki/tasks.py` (line 288). Each page from the `WikiStructureModel` is processed by `_generate_page()` (line 369), which builds a per-page prompt via `build_page_prompt()` in `prompts.py` (line 25). The prompt includes the page title and a pre-built Markdown list of relevant file links with host-specific URLs (GitHub `blob/{branch}`, GitLab `-/blob/{branch}`, Bitbucket `src/{branch}`). The prompt instructs the model to: start with a `<details>" summary="Relevant source files"` block citing >=5 files, use Mermaid diagrams extensively (flowcharts, sequence diagrams, class diagrams — all top-down via `graph TD`), create markdown tables, and cite sources using empty-parenthesis citations like `Sources: [path/to/file.py:10-25]()` (line 108-118).

Generation streams the LLM response through the `research_chat()` pipeline (same as RAG chat), then post-processes it in `post_process_wiki_content()` (`api/services/wiki/content.py`, line 84). This function: (1) rebuilds the `<details>` block with real hyperlinks from the known file list, (2) resolves all empty-parenthesis citations into proper GitHub/GitLab/Bitbucket URLs with line anchors (`#L10-L25`), (3) handles generic file-path citations and basename-only references, and (4) strips redundant empty parens.

Pages are generated with configurable concurrency: `WIKI_PAGE_CONCURRENCY` defaults to 1 (sequential) but each page gets up to `WIKI_PAGE_RETRIES + 1` attempts before falling back to an error-placeholder page (line 268-285). Once all pages are done, the wiki is saved as a JSON cache file (`deepwiki_cache_{type}_{owner}_{repo}_{lang}.json`) in the wiki cache directory via `save_wiki_cache()` (`io.py`, line 52). Existing cache is detected at submission time to skip regeneration entirely (tasks.py line 176).


Citations: [api/services/wiki/prompts.py:25-132](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/prompts.py#L25-L132) · [api/services/wiki/content.py:57-151](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/content.py#L57-L151) · [api/services/wiki/io.py:52-69](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/io.py#L52-L69) · [api/services/wiki/content.py:23-36](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/content.py#L23-L36)

### How are model providers configured? (answered)

Model providers are configured through JSON config files and a client registry pattern. The primary config is `api/config/generator.json`, which defines providers and their models. The default provider is `google` with `gemini-2.5-flash` as the default model; OpenAI (`gpt-5-nano` as default), OpenRouter, Ollama, Bedrock, Azure, DashScope, and LiteLLM are all supported (generator.json lines 4-198). A LiteLLM variant of the config lives in `generator.litellm.json` with `litellm` as the default provider.

Each provider has a corresponding `ChatStreamer` subclass in `api/chat/_stream.py` that registers itself via the `provider` class attribute: `OllamaChatStreamer`, `OpenRouterChatStreamer`, `OpenAIChatStreamer`, `AzureChatStreamer`, `LiteLLMChatStreamer`, `BedrockChatStreamer`, `DashScopeChatStreamer`, `GoogleGenerativeChatStreamer`, and `AnthropicChatStreamer`. The `ChatStreamer.create()` factory method (line 40) looks up the provider name in a registry dict and instantiates the right class.

For per-stage model selection, `get_model_config()` in `config.py` (line 364) looks up a provider's configuration, resolves the model client class, and extracts model-specific parameters (temperature, top_p, etc.). The same provider+model tuple is passed through all requests (`ChatCompletionRequest` and `WikiTaskRequest` via `RepoRequestBase` schema, `api/schemas/base.py` line 9). The embedder is configured separately via `DEEPWIKI_EMBEDDER_TYPE` env var (openai | google | ollama | bedrock) and its own config file (`embedder.json`).

OpenAI-compatible endpoints are supported via several paths. The `LiteLLMClient` (`api/clients/litellm.py`) extends the OpenAI client, connecting to a LiteLLM proxy at `LITELLM_BASE_URL` (default localhost:4000). The `OpenAIClient` uses the standard OpenAI API. Any OpenAI-compatible endpoint can be configured by pointing the client to a different base URL. Local models run through `OllamaClient` (ollama.py) or via the Ollama provider with local models like `qwen3:1.7b`, `llama3:8b`, and `qwen3:8b` (generator.json lines 116-141).


Citations: [api/config/generator.json:1-198](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config/generator.json#L1-L198) · [api/config.py:364-417](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config.py#L364-L417) · [api/clients/litellm.py:8-69](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/clients/litellm.py#L8-L69) · [api/schemas/base.py:9-44](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/schemas/base.py#L9-L44) · [api/config/generator.litellm.json:1-30](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config/generator.litellm.json#L1-L30)

### How is interactive Q&A / chat implemented? (answered)

Interactive Q&A over a repository is implemented through a WebSocket-based chat system. The frontend's `Ask` component (`src/components/Ask.tsx`) provides the UI, with three modes: Fast, Deep Research, and Codemap. When a user submits a question, the frontend opens a WebSocket to `/ws/chat` (`src/utils/websocketClient.ts`, line 64), sends a `ChatCompletionRequest` as JSON, and receives streaming text chunks. On WebSocket failure, it falls back to HTTP POST at `/api/chat/stream` (`src/app/api/chat/stream/route.ts`).

The backend endpoint is `handle_websocket_chat()` in `api/routers/chat.py` (line 20). It delegates to `research_chat()` (`api/services/research.py`), which: (1) prepares or loads the FAISS RAG index for the repo, (2) retrieves relevant document chunks, (3) loads conversation history from the `Memory` class into XML-style `<turn>` blocks, (4) picks the appropriate iteration prompt, and (5) streams the LLM response via `ChatStreamer.respond_stream()`. If the prompt exceeds 7500 tokens, RAG context is dropped and the model answers from its training data alone (line 146-148).

Deep Research mode (`api/prompts.py` lines 60-151) runs up to 5 automated iterations: iteration 1 produces a research plan, iterations 2-4 investigate deeper aspects, and iteration 5 produces a final conclusion. Each iteration builds on the previous conversation history. The frontend's `Ask.tsx` auto-continues through iterations (line 620-637), extracting research stages from the response (`## Research Plan`, `## Research Update N`, `## Final Conclusion`) and displaying a stage navigator. A max of 5 iterations is hard-coded (research.py line 219), with force-complete logic on the frontend (Ask.tsx line 501-508).

Codemap mode (`api/services/codemap.py`) is a separate multi-phase pipeline: RAG retrieval → two-step JSON generation (skeleton + enrich with Mermaid diagrams) → citation grounding against real source files → NDJSON streamed over `/ws/codemap`. Results are displayed as structured step-by-step guides with a side-by-side CodeViewer for cited source files.

Conversation history is maintained client-side in the `conversationTurns` array (Ask.tsx line 92), with model/provider selection exposed through a modal. Unlike the wiki generation pipeline, Q&A does not cache results on the server.


Citations: [api/services/research.py:60-311](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/research.py#L60-L311) · [src/components/Ask.tsx:110-144](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/src/components/Ask.tsx#L110-L144) · [src/utils/websocketClient.ts:64-96](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/src/utils/websocketClient.ts#L64-L96) · [api/routers/chat.py:20-66](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/routers/chat.py#L20-L66) · [api/prompts.py:60-151](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/prompts.py#L60-L151) · [api/services/codemap.py:211-298](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/codemap.py#L211-L298)
