LLMs Technical Reviews

How are individual pages generated?

Per-page prompts; source-file citations; diagrams (Mermaid); parallelism; caching and regeneration.

Verdict

deepwiki-open builds each page with one LLM call. build_page_prompt() passes the title and the page’s file links. The model must open with a <details> source list, use Mermaid diagrams, and cite as Sources: [path:10-25](). Context is the top-20 retrieved chunks. post_process_wiki_content() then turns those citations into GitHub, GitLab or Bitbucket links with line anchors. Pages run with DEEPWIKI_WIKI_PAGE_CONCURRENCY (default 1) and two retries, then an error placeholder. The whole wiki is saved as one JSON cache file. A cached wiki is never updated, only deleted and rebuilt.

OpenDeepWiki runs an agent per page with the content-generator prompt. It has 6 source-tool calls and writes with WriteDoc plus up to 4 AppendDoc calls, so a page can be longer than one response. Links use a File Reference Base URL, and DocTool stores the files actually read as SourceFiles. The prompt includes detailed Mermaid syntax rules, and two normalizers clean the output. The first page runs alone to warm the prompt cache, then 5 pages run in parallel, each with a 60-minute timeout. Pages that already exist are skipped, and new commits trigger IncrementalUpdateAsync() on the affected pages only.

deepwiki-open is simpler and has stronger line-level citation repair. OpenDeepWiki is better for long pages, large repos and wikis that must follow the code over time.

Per-project answers

AsyncFuncAI/deepwiki-open

answered

Individual wiki pages are generated by _generate_pages() in api/services/wiki/tasks.py (line 288). Each page from the WikiStructureModel is processed by _generate_page() (line 369), which builds a per-page prompt via build_page_prompt() in prompts.py (line 25). The prompt includes the page title and a pre-built Markdown list of relevant file links with host-specific URLs (GitHub blob/{branch}, GitLab -/blob/{branch}, Bitbucket src/{branch}). The prompt instructs the model to: start with a <details>" summary="Relevant source files" block citing >=5 files, use Mermaid diagrams extensively (flowcharts, sequence diagrams, class diagrams — all top-down via graph TD), create markdown tables, and cite sources using empty-parenthesis citations like Sources: [path/to/file.py:10-25]() (line 108-118).

Generation streams the LLM response through the research_chat() pipeline (same as RAG chat), then post-processes it in post_process_wiki_content() (api/services/wiki/content.py, line 84). This function: (1) rebuilds the <details> block with real hyperlinks from the known file list, (2) resolves all empty-parenthesis citations into proper GitHub/GitLab/Bitbucket URLs with line anchors (#L10-L25), (3) handles generic file-path citations and basename-only references, and (4) strips redundant empty parens.

Pages are generated with configurable concurrency: WIKI_PAGE_CONCURRENCY defaults to 1 (sequential) but each page gets up to WIKI_PAGE_RETRIES + 1 attempts before falling back to an error-placeholder page (line 268-285). Once all pages are done, the wiki is saved as a JSON cache file (deepwiki_cache_{type}_{owner}_{repo}_{lang}.json) in the wiki cache directory via save_wiki_cache() (io.py, line 52). Existing cache is detected at submission time to skip regeneration entirely (tasks.py line 176).

AIDotNet/OpenDeepWiki

answered

Per-page prompts. Each wiki page is generated by an LLM agent using the content-generator prompt (src/OpenDeepWiki/prompts/content-generator.md). WikiGenerator.GenerateDocumentsAsync() (src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs:366) iterates over all leaf catalog items. For each item, GenerateDocumentContentAsync() (line 796) builds a user message with the catalog path/title, a git-based File Reference Base URL, and budget limits. The agent receives a GitTool (capped at MaxDocumentSourceToolCalls, default 6), a DocTool (for WriteDoc/AppendDoc/EditDoc), and a tool budget (MaxDocumentToolCalls, default 14; MaxDocumentAppendOperations, default 4).

Source-file citations. The agent is instructed to read 2-4 key source files, extract code snippets, and embed them with Markdown blockquote source links using the runtime File Reference Base URL (line 843–848). The DocTool automatically tracks which files were read; these are stored as a JSON array in DocFile.SourceFiles on the entity.

Mermaid diagrams. The prompt mandates at least one Mermaid diagram per page (architecture/flowchart/sequence/class/ER). Extensive Mermaid syntax rules are embedded in the prompt (line 850–900), including node ID naming, subgraph sg_ prefix requirements, and edge label quoting — reflecting real renderer constraints encountered in production.

Parallelism. GenerateDocumentsAsync() uses Parallel.ForEachAsync with MaxDegreeOfParallelism from WikiGeneratorOptions.ParallelCount (default 5, line 60). Before parallel execution, it "warms" the prompt cache by generating the first document sequentially (line 544–559), then processes the remaining items in parallel. Timed-out or failed documents are logged but don't stop the batch (line 497–542) unless ThrowOnPartialFailure is set. Each document has a timeout of DocumentGenerationTimeoutMinutes (default 60).

Caching and regeneration. Documents that already have persisted content via DocFileId are skipped (line 414–421). RegenerateDocumentAsync() (line 611) allows forced regeneration of a single path. Incremental updates (via IncrementalUpdateAsync, line 660) use the incremental-updater prompt with a restricted toolset (no WriteCatalog) to edit only affected documents based on git diff results between commits.

← How is the wiki structure (table of contents) determined? · How are model providers configured? →