AsyncFuncAI/deepwiki-open
FastAPI service and Next.js UI that embed a repo into a FAISS index, then generate a cited, Mermaid-heavy wiki with one prompt per page.
Overview
deepwiki-open is a self-hosted wiki generator built as a classic RAG pipeline. A Python FastAPI backend (api/) clones or reads a repository, splits every eligible file into word chunks, embeds them, and saves a FAISS-ready index as a pickle. It then asks an LLM for a wiki outline in XML and writes each page with a separate LLM call. A Next.js frontend (src/) renders the wiki and offers chat over the same index.
The LLM never reads files on its own. Everything it sees comes from the prompt: the file tree and README for the outline, then the top-20 retrieved chunks for each page or chat turn. This keeps the design simple and cheap to run. It also means page quality depends on what the retriever happens to return, not on what the model would have chosen to read.
It fits a single team that wants a local DeepWiki for GitHub, GitLab, Bitbucket or local repositories, with a free choice of model provider (Google, OpenAI, OpenRouter, Ollama, Bedrock, Azure, DashScope, LiteLLM, Anthropic). It has no user accounts, no database and no incremental update. The output is one JSON cache file per repository and language.
Architecture
flowchart TD
UI["Next.js UI (src/)"] -->|"POST /wiki/tasks"| WR["routers/wiki.py"]
UI -->|"WS /ws/chat"| CR["routers/chat.py"]
WR --> TR["TaskRegistry (tasks.py)"]
TR --> GEN["generate_repo_wiki()"]
GEN --> REPO["Repo: clone or local path"]
GEN --> IDX["prepare_repo_index()"]
IDX --> PIPE["split + embed (rag/pipeline.py)"]
PIPE --> PKL[("LocalDB .pkl")]
GEN --> STRUCT["_determine_structure()"]
GEN --> PAGES["_generate_pages()"]
STRUCT --> RC["research_chat()"]
PAGES --> RC
CR --> RC
RC --> FAISS["FAISSRetriever top_k=20"]
FAISS --> PKL
RC --> CS["ChatStreamer.create(provider)"]
CS --> LLM["LLM provider"]
PAGES --> CACHE[("wikicache JSON")]
| Component | Path | Role |
|---|---|---|
| App entry | api/main.py | Mounts the system, auth, repo, wiki, chat and codemap routers |
| Wiki task API | api/routers/wiki.py | Submit, list, poll and SSE-stream generation tasks |
| Task runner | api/services/wiki/tasks.py | In-memory registry, semaphore, state machine |
| Repository access | api/repository.py | Shallow clone with token injection, or use a local path in place |
| File filter | api/config.py | iterate_files(): extension allow-list plus include/exclude rules |
| Indexing | api/rag/pipeline.py | Line-tracking splitter, embedder, LocalDB pickle |
| Retrieval + prompt | api/services/research.py | research_chat(): one path for outline, pages and chat |
| Providers | api/chat/_stream.py | ChatStreamer registry, one subclass per provider |
| Prompts | api/services/wiki/prompts.py | Outline (XML) and page prompts |
| Post-processing | api/services/wiki/content.py | Turns [path:10-25]() citations into real host links |
How a request flows
- The client calls
POST /wiki/tasks.submit_wiki_taskhands aWikiTasktoTaskRegistry.submit(). If a task for the same repo is active, the caller joins it. If a cache file already exists for that repo and language, the call returnsfrom_cacheand nothing is regenerated (tasks.py#L161-L190). TaskRegistry._run()waits on a semaphore sized byDEEPWIKI_MAX_CONCURRENT_WIKI_TASKS, then runsgenerate_repo_wiki().- Indexing. If no
.pklexists,prepare_repo_index()builds aRAGobject and callsaprepare_retriever().Repo.download()does a--depth=1 --single-branchclone; local paths are used as they are (repository.py#L191-L222). Files come fromiterate_files(). Files aboveMAX_EMBEDDING_TOKENS * 10tokens are skipped (pipeline.py#L112-L130). The rest are split, embedded and saved withLocalDB.save_state(). - Outline.
_determine_structure()reads the file list and the shortest-path README, and builds the outline prompt withbuild_structure_prompt(). It sends the prompt throughresearch_chat()and parses the reply withparse_wiki_structure(). Any failure here fails the whole task. - Pages.
_generate_pages()runs_generate_page()for every page under a second semaphore (DEEPWIKI_WIKI_PAGE_CONCURRENCY, default 1). Each page is retriedWIKI_PAGE_RETRIEStimes, then replaced by an error placeholder, so one bad page does not fail the wiki. - Inside
research_chat(), the last user message is the retrieval query. The top-20 chunks are grouped by file and annotated with[lines a-b].prompt_builder()wraps them in<START_OF_CONTEXT>tags (_prompts.py#L6-L31). ChatStreamer.create()picks the provider class and streams the answer.post_process_wiki_content()rebuilds the<details>source list and resolves citations.save_wiki_cache()writes the JSON file. Progress is visible throughGET /wiki/tasks/{id}/stream(SSE, one event per second).
Key components
Retrieval index
The splitter config is split_by: word, chunk_size: 350, chunk_overlap: 100, and the retriever uses top_k: 20 (api/config/embedder.json). LineTrackingTextSplitter finds each chunk back in its parent text and stores 1-based start_line/end_line, so the model can cite real line numbers (pipeline.py#L174-L208). The index is a pickle at <root>/databases/<repo>.pkl, and its existence is the only “already indexed” check (pipeline.py#L163-L171). A FAISSRetriever is built in memory from the stored vectors (rag.py#L288-L296).
One prompt path for everything
The outline, every page, and every chat turn all go through research_chat(). Two details matter. First, the retrieval query for a page is the whole page prompt, which is mostly fixed instructions plus the title and file links. Second, if that last message is over MAX_INPUT_TOKENS = 7500 tokens, retrieval is skipped (research.py#L22-L23, #L64-L73). For most real repos the outline prompt carries the full file list and README, so it crosses that limit and the outline is built from the tree and README only.
Page prompt and citations
build_page_prompt() tells the model it has “the full content” of the listed files and must cite at least five of them. In fact it only receives retrieved chunks. Citations are written as Sources: [path:10-25]() with empty parentheses (prompts.py#L108-L118). post_process_wiki_content() resolves them against the page’s file list, then generic paths, then bare file names, into GitHub, GitLab or Bitbucket URLs.
Providers
Each ChatStreamer subclass registers itself by its provider attribute (_stream.py#L29-L48). get_model_config() reads generator.json and falls back to the provider’s default model’s parameters for unknown model names (config.py#L364-L417). The OpenRouter client always sends stream: False and sets a 60 s timeout, even though the streamer asks for streaming (openrouter.py#L140-L157).
Chat, Deep Research and Codemap
/ws/chat (chat.py#L20-L66) streams research_chat() output as text frames. History is sent by the client on each request and replayed as <turn> blocks. Deep Research is a prompt switch on research_iteration (plan, updates, final answer at iteration 5); the frontend drives the loop. Codemap (api/services/codemap.py) is a separate two-call pipeline that returns NDJSON.
Extending it
- New provider: subclass
ChatStreamerwith aprovidername and add a block inapi/config/generator.json. Registration is automatic. - Embedder:
DEEPWIKI_EMBEDDER_TYPE(openai, google, ollama, bedrock) selects a block inembedder.json. Any OpenAI-compatible embedding endpoint works throughOPENAI_BASE_URL. - Filters:
api/config/repo.jsonholds the extension lists and default excluded dirs. Requests can addexcluded_dirs,excluded_files,included_dirs,included_files. - Prompts: outline and page prompts are plain f-strings in
api/services/wiki/prompts.py; chat and research prompts are inapi/prompts.py.
Running it
The repo ships a docker-compose.yml that runs the API on 8001 and the UI on 3000 and mounts ~/.adalflow for clones and indexes. You need one generator key (Google by default) and one embedder key; the default embedder is OpenAI text-embedding-3-small.
This site ran the API headless on every repository in our catalog. Notes from that run:
- We drove it with
POST /wiki/tasksand a local repo path, read-only mounted, so it never cloned from GitHub. - Without an OpenAI key you can point the OpenAI embedder at any OpenAI-compatible endpoint through
OPENAI_BASE_URL(we usedopenai/text-embedding-3-smallat 256 dims) and a replacedembedder.json. - The OpenRouter client’s 60 s non-streaming timeout killed whole-repo outline calls. A failed call comes back as error text, the XML parse fails, and the whole task fails. We patched the timeout to 900 s.
MAX_CONCURRENT_WIKI_TASKSdefaults toos.cpu_count() // 2(tasks.py#L61-L66). Python reports host cores, not the container CPU limit, so on a large host it would launch far more tasks than a small container can run; set it explicitly.- Tasks live only in memory. A restart loses running work, and finished tasks drop out of the registry after 300 s; after that, only the cache file shows the result.
Measured on the same three repos, with DeepSeek V4 Flash for both tools and 3 pages in parallel here: llm-scraper 11 pages in 285 s, OpenBot 12 pages in 798 s, rakazo 10 pages in 512 s. deepwiki-open records no token usage, so we have no per-wiki cost for it. The OpenDeepWiki page has its numbers.
Strengths and caveats
- Strength: simple and portable. One Python process, files on disk, no database. Nine provider classes and Ollama make fully local runs possible.
- Strength: good citation plumbing. Line-tracked chunks plus
post_process_wiki_content()give pages real#Lx-Lylinks. - Strength: fixed, predictable cost. One outline call plus one call per page, with a bounded prompt.
- Caveat: retrieval-bound pages. The model sees only 20 chunks per page, retrieved with the page prompt as the query, even though the prompt claims full file access.
- Caveat: outline skips RAG on large repos because of the 7500-token check on the query, and the file list is sent as a Python list literal.
- Caveat: argument-order bug.
_determine_structure()passes(excluded_dirs, excluded_files, included_dirs, included_files)by position, butread_repo_file_tree()expects(included_files, included_dirs, excluded_files, excluded_dirs)(tasks.py#L328-L335). A request with exclusions therefore builds the outline in inclusion mode. The index itself uses keyword arguments and is not affected. - Caveat: no incremental update. A cached wiki is returned as-is; to refresh it you delete the cache and regenerate everything.
- Caveat: no output check.
stream_and_fallback()yields provider errors as text, so a failed page can be saved as an error string instead of being retried. Nothing checks the page shape either: when we ran deepwiki-open on its own repo, the “Project Overview” page came back as about 600 lines of the model’s reasoning, and it was saved as a finished page. - Caveat: leftover frontend paths. The UI still contains a
POST /api/wiki_cachecall for saving a finished wiki, but the backend only definesGETandDELETEfor that route (wiki.py#L139-L165). The backend task writes the cache itself.
Sources: code at d92819a, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (not available), verified Q&A.
How it answers the Open-source DeepWiki questions
Each answer was drafted by a code-reading agent at commit d92819a. Its citations were checked mechanically. Compare with the other open-source deepwiki →
How is a repository ingested and chunked?
answeredA repository is ingested via the Repo class in api/repository.py (line 140). It supports GitHub, GitLab, Bitbucket, and local paths — each with its own _clone_from_* function that handles access tokens via URL injection (oauth2:token@ for GitLab, token@ for GitHub, x-bitbucket-api-token-auth for Bitbucket). Clones use git.Repo.clone_from with --depth=1 --single-branch for shallow clones. Local paths skip cloning entirely and are used in-place. The cloned repo path is ~/.deepwiki/repo/.
File filtering is centralized in api/config.py's iterate_files() (line 467). It walks the repo directory and applies: (1) extension whitelisting from code_extensions and doc_extensions config arrays, (2) exclusion mode using default file_filters.excluded_dirs (like node_modules, .venv, .git, dist) merged with request-level exclusions, or (3) inclusion mode when included_dirs/included_files are specified. Default excluded dirs are extensive (line 3-40 of repo.json).
Chunking happens in the RAG pipeline, not at ingestion. api/rag/pipeline.py's LineTrackingTextSplitter (line 174) uses adalflow's TextSplitter with split_by: word, chunk_size: 350, chunk_overlap: 100 (from embedder.json line 36-41). Each chunk gets annotated with its 1-based start_line/end_line within the original file. Files exceeding MAX_EMBEDDING_TOKENS * 10 (82K tokens) are skipped (line 126). The pipeline then embeds each chunk and saves to a local /.adalflow/databases/LocalDB pickle file at `(line 247-274). Large-repo limits are configurable viaMAX_CONCURRENT_WIKI_TASKS(defaults to half CPU cores) andWIKI_PAGE_CONCURRENCY(default 1) inapi/services/wiki/tasks.py` (lines 62-68).
How is retrieval (RAG) implemented?
answeredRetrieval is implemented via a FAISS-based RAG pipeline using the adalflow framework. The core class is RAG in api/rag/rag.py (line 134), which wraps a FAISSRetriever (line 290) configured with top_k: 20 (from embedder.json). Documents are embedded using configurable embedders: OpenAI text-embedding-3-small (256 dims) by default, Google gemini-embedding-001, Ollama nomic-embed-text, or Bedrock amazon.titan-embed-text-v2:0 — selected via the DEEPWIKI_EMBEDDER_TYPE env var (config.py line 65).
The retrieval pipeline in research_chat() (api/services/research.py, line 60) first prepares the retriever index (loading from a .pkl pickle file or creating it fresh), then calls rag.acall(query) which invokes the FAISS retriever. Retrieved documents are grouped by file path and each chunk is annotated with its line range (line 178-184). The grouped context is assembled into a structured prompt that includes file-path headers and line-numbered chunks inside <START_OF_CONTEXT> / <END_OF_CONTEXT> tags (line 146-192).
The context is then injected into prompt_builder() (api/chat/_prompts.py, line 6), which builds a final prompt combining the system prompt, conversation history, context block, and user query. The response streamer (ChatStreamer in _stream.py) sends the full prompt to the LLM and streams back chunks. If the input exceeds ~7500 tokens, RAG context is omitted and a fallback note is inserted instead (line 27-29 of _prompts.py). The Memory class in rag.py (line 79) accumulates conversation turns for multi-turn dialogue.
How is the wiki structure (table of contents) determined?
answeredThe wiki structure (table of contents) is determined by an LLM call. In api/services/wiki/tasks.py, _determine_structure() (line 315) orchestrates the process. First it reads the cloned repo's file tree and README via read_repo_file_tree() in structure.py (line 20), which returns sorted file paths filtered through iterate_files() plus the README content. The repo's default branch is detected with git rev-parse --abbrev-ref HEAD (line 51).
These inputs are fed into build_structure_prompt() in prompts.py (line 213). The prompt constructs an XML schema asking the LLM to output a <wiki_structure> block containing a title, description, and individual <page> elements — each with an id, title, importance (high/medium/low), related pages, and a list of <file_path> citations. For "comprehensive" wikis, it also includes <sections> grouping pages and supporting nested subsections (line 135-185); "concise" omits sections and targets 4-6 pages instead of 8-12 (line 187-209).
The LLM's raw XML response is parsed by parse_wiki_structure() in structure.py (line 179). This function is robust: it strips markdown fences, escapes bare & characters, handles truncated responses (missing </wiki_structure>) by salvaging complete blocks, and falls back to regex-based extraction if strict xml.etree.ElementTree parsing fails. The result is a WikiStructureModel containing WikiPage and WikiSection objects (schemas in api/schemas/wiki.py).
How are individual pages generated?
answeredIndividual wiki pages are generated by _generate_pages() in api/services/wiki/tasks.py (line 288). Each page from the WikiStructureModel is processed by _generate_page() (line 369), which builds a per-page prompt via build_page_prompt() in prompts.py (line 25). The prompt includes the page title and a pre-built Markdown list of relevant file links with host-specific URLs (GitHub blob/{branch}, GitLab -/blob/{branch}, Bitbucket src/{branch}). The prompt instructs the model to: start with a <details>" summary="Relevant source files" block citing >=5 files, use Mermaid diagrams extensively (flowcharts, sequence diagrams, class diagrams — all top-down via graph TD), create markdown tables, and cite sources using empty-parenthesis citations like Sources: [path/to/file.py:10-25]() (line 108-118).
Generation streams the LLM response through the research_chat() pipeline (same as RAG chat), then post-processes it in post_process_wiki_content() (api/services/wiki/content.py, line 84). This function: (1) rebuilds the <details> block with real hyperlinks from the known file list, (2) resolves all empty-parenthesis citations into proper GitHub/GitLab/Bitbucket URLs with line anchors (#L10-L25), (3) handles generic file-path citations and basename-only references, and (4) strips redundant empty parens.
Pages are generated with configurable concurrency: WIKI_PAGE_CONCURRENCY defaults to 1 (sequential) but each page gets up to WIKI_PAGE_RETRIES + 1 attempts before falling back to an error-placeholder page (line 268-285). Once all pages are done, the wiki is saved as a JSON cache file (deepwiki_cache_{type}_{owner}_{repo}_{lang}.json) in the wiki cache directory via save_wiki_cache() (io.py, line 52). Existing cache is detected at submission time to skip regeneration entirely (tasks.py line 176).
How are model providers configured?
answeredModel providers are configured through JSON config files and a client registry pattern. The primary config is api/config/generator.json, which defines providers and their models. The default provider is google with gemini-2.5-flash as the default model; OpenAI (gpt-5-nano as default), OpenRouter, Ollama, Bedrock, Azure, DashScope, and LiteLLM are all supported (generator.json lines 4-198). A LiteLLM variant of the config lives in generator.litellm.json with litellm as the default provider.
Each provider has a corresponding ChatStreamer subclass in api/chat/_stream.py that registers itself via the provider class attribute: OllamaChatStreamer, OpenRouterChatStreamer, OpenAIChatStreamer, AzureChatStreamer, LiteLLMChatStreamer, BedrockChatStreamer, DashScopeChatStreamer, GoogleGenerativeChatStreamer, and AnthropicChatStreamer. The ChatStreamer.create() factory method (line 40) looks up the provider name in a registry dict and instantiates the right class.
For per-stage model selection, get_model_config() in config.py (line 364) looks up a provider's configuration, resolves the model client class, and extracts model-specific parameters (temperature, top_p, etc.). The same provider+model tuple is passed through all requests (ChatCompletionRequest and WikiTaskRequest via RepoRequestBase schema, api/schemas/base.py line 9). The embedder is configured separately via DEEPWIKI_EMBEDDER_TYPE env var (openai | google | ollama | bedrock) and its own config file (embedder.json).
OpenAI-compatible endpoints are supported via several paths. The LiteLLMClient (api/clients/litellm.py) extends the OpenAI client, connecting to a LiteLLM proxy at LITELLM_BASE_URL (default localhost:4000). The OpenAIClient uses the standard OpenAI API. Any OpenAI-compatible endpoint can be configured by pointing the client to a different base URL. Local models run through OllamaClient (ollama.py) or via the Ollama provider with local models like qwen3:1.7b, llama3:8b, and qwen3:8b (generator.json lines 116-141).
How is interactive Q&A / chat implemented?
answeredInteractive Q&A over a repository is implemented through a WebSocket-based chat system. The frontend's Ask component (src/components/Ask.tsx) provides the UI, with three modes: Fast, Deep Research, and Codemap. When a user submits a question, the frontend opens a WebSocket to /ws/chat (src/utils/websocketClient.ts, line 64), sends a ChatCompletionRequest as JSON, and receives streaming text chunks. On WebSocket failure, it falls back to HTTP POST at /api/chat/stream (src/app/api/chat/stream/route.ts).
The backend endpoint is handle_websocket_chat() in api/routers/chat.py (line 20). It delegates to research_chat() (api/services/research.py), which: (1) prepares or loads the FAISS RAG index for the repo, (2) retrieves relevant document chunks, (3) loads conversation history from the Memory class into XML-style <turn> blocks, (4) picks the appropriate iteration prompt, and (5) streams the LLM response via ChatStreamer.respond_stream(). If the prompt exceeds 7500 tokens, RAG context is dropped and the model answers from its training data alone (line 146-148).
Deep Research mode (api/prompts.py lines 60-151) runs up to 5 automated iterations: iteration 1 produces a research plan, iterations 2-4 investigate deeper aspects, and iteration 5 produces a final conclusion. Each iteration builds on the previous conversation history. The frontend's Ask.tsx auto-continues through iterations (line 620-637), extracting research stages from the response (## Research Plan, ## Research Update N, ## Final Conclusion) and displaying a stage navigator. A max of 5 iterations is hard-coded (research.py line 219), with force-complete logic on the frontend (Ask.tsx line 501-508).
Codemap mode (api/services/codemap.py) is a separate multi-phase pipeline: RAG retrieval → two-step JSON generation (skeleton + enrich with Mermaid diagrams) → citation grounding against real source files → NDJSON streamed over /ws/codemap. Results are displayed as structured step-by-step guides with a side-by-side CodeViewer for cited source files.
Conversation history is maintained client-side in the conversationTurns array (Ask.tsx line 92), with model/provider selection exposed through a modal. Unlike the wiki generation pipeline, Q&A does not cache results on the server.