How is the wiki structure (table of contents) determined?
Prompt/agent that proposes sections and pages; inputs used (file tree, README); output format.
Verdict
Both tools ask an LLM for the table of contents. They differ in what the model sees and how the answer comes back.
deepwiki-open makes one plain completion call. build_structure_prompt() includes the complete filtered file list and the README, and asks for <wiki_structure> XML with 4 to 6 pages (concise) or 8 to 12 pages in sections (comprehensive). Each page lists its relevant files. parse_wiki_structure() tolerates code fences, stray & and truncated output. The model cannot check any file, and in this version an argument-order bug turns request exclusions into inclusions for this step.
OpenDeepWiki runs an agent with the catalog-generator prompt. Its input is a TOON-encoded tree limited by the scan plan, the README (up to 4,000 chars), key files and entry points. It can make up to 4 confirmation calls (ListFiles, Grep, ReadFile) and then saves a nested JSON catalog with WriteCatalog. Page count is not fixed. The prompt asks for “right-sized” coverage, so big repos get many more pages (61 for OpenBot in our runs, against 12 from deepwiki-open).
deepwiki-open gives a predictable, small wiki at a known cost. OpenDeepWiki gives a catalog that grows with the codebase and is checked against a few real files.
Per-project answers
AsyncFuncAI/deepwiki-open
answeredThe wiki structure (table of contents) is determined by an LLM call. In api/services/wiki/tasks.py, _determine_structure() (line 315) orchestrates the process. First it reads the cloned repo's file tree and README via read_repo_file_tree() in structure.py (line 20), which returns sorted file paths filtered through iterate_files() plus the README content. The repo's default branch is detected with git rev-parse --abbrev-ref HEAD (line 51).
These inputs are fed into build_structure_prompt() in prompts.py (line 213). The prompt constructs an XML schema asking the LLM to output a <wiki_structure> block containing a title, description, and individual <page> elements — each with an id, title, importance (high/medium/low), related pages, and a list of <file_path> citations. For "comprehensive" wikis, it also includes <sections> grouping pages and supporting nested subsections (line 135-185); "concise" omits sections and targets 4-6 pages instead of 8-12 (line 187-209).
The LLM's raw XML response is parsed by parse_wiki_structure() in structure.py (line 179). This function is robust: it strips markdown fences, escapes bare & characters, handles truncated responses (missing </wiki_structure>) by salvaging complete blocks, and falls back to regex-based extraction if strict xml.etree.ElementTree parsing fails. The result is a WikiStructureModel containing WikiPage and WikiSection objects (schemas in api/schemas/wiki.py).
AIDotNet/OpenDeepWiki
answeredThe wiki structure (table of contents) is generated by an LLM agent using the catalog-generator prompt (src/OpenDeepWiki/prompts/catalog-generator.md). WikiGenerator.GenerateCatalogAsync() (src/OpenDeepWiki/Services/Wiki/WikiGenerator.cs:245) is the entry point. It first calls CollectRepositoryContextAsync() to gather the project's directory tree (TOON format), README content, key files, and entry points. Then it loads the catalog-generator prompt template, creates a GitTool (for source exploration), a CatalogTool (wrapping CatalogStorage to write the catalog), and a DocumentSourceToolBudget that limits source tool calls to the configured MaxCatalogSourceToolCalls (default 4). The AI agent is instructed to design a DeepWiki-style information architecture — top-level topic domains with independently readable deep-dive leaf pages — not a file/class listing.
Inputs. The agent receives a runtime context containing the directory tree, README content, key files, entry points, and project type. It uses the tools sparingly for confirmation reads/greps before writing the catalog.
Output format. The catalog is JSON via the WriteCatalog tool (CatalogTool.WriteAsync), following:
{"items": [{"title": "...", "path": "lowercase-hyphen-path", "order": N, "children": []}]}
Each leaf node gets a DocFileId pointing to generated content. Parent nodes are navigation groupings only.
Design rules. The prompt (line 25–75) emphasizes right-sized coverage (no fixed counts), no over-compression (many independent systems must not be hidden inside few chapters), verifying via source tools first, and never fabricating entries. Page titles should be in the runtime target language. The CatalogStorage class (src/OpenDeepWiki/Services/Wiki/CatalogStorage.cs:12) persists the catalog to DocCatalog entities in the database, keyed by BranchLanguageId. Parent-child relationships are tracked via ParentId and an Order field.
← How is retrieval (RAG) implemented? · How are individual pages generated? →