he-yufeng/RepoWiki
Python CLI and FastAPI app that builds a fixed-shape repo wiki from four LLM passes, an import-graph PageRank and TF-IDF chat.
Overview
RepoWiki is a small Python tool that turns a local directory or a git URL into a browsable wiki. It ships as a repowiki CLI (Click and Rich) with an optional FastAPI server and a React front end, installed through the web extra. The design is deliberately cheap. One file scan feeds two zero-LLM analyses: a regex import graph ranked with PageRank, and a TF-IDF index for chat. Then four kinds of LLM call fill a wiki whose page layout is fixed in code.
That fixed layout is the main difference from DeepWiki-style tools that let a model plan the table of contents. RepoWiki always produces the same skeleton: Overview, Architecture, Knowledge Cards, one page per top-level module, Reading Guide, Dependencies and Symbol Index. The LLM only fills content slots with JSON. Modules are not discovered semantically. They are the first directory under the root, or the second one under src/, lib/, pkg/, internal/ or app/.
Two commands use no model at all: repowiki map prints the PageRank file ranking, and repowiki diff <refspec> orders a diff’s changed files by importance and blast radius for review. Both are aimed at coding agents that need a reading order.
Architecture
flowchart LR
IN["CLI or POST /api/scan"] --> ING["ingest_local / ingest_github"]
ING --> SCAN["scan_directory"]
SCAN --> PC["ProjectContext"]
PC --> AN["Analyzer: 4 LLM passes"]
AN --> LLM["LLMClient (litellm)"]
AN <--> CACHE["SQLite cache"]
PC --> DG["DependencyGraph + PageRank"]
DG --> AN
AN --> WB["WikiBuilder"]
DG --> WB
WB --> EXP["Markdown / JSON / HTML export"]
WB --> API["FastAPI + React UI"]
PC --> RAG["SimpleRAG TF-IDF index"]
RAG --> CHAT["repowiki chat / chat SSE"]
CHAT --> LLM
| Component | Path | Role |
|---|---|---|
| CLI | src/repowiki/cli.py |
scan, chat, serve, map, diff, config, cache-clear |
| Ingest | src/repowiki/ingest/ |
Shallow-clones GitHub/GitLab/Bitbucket URLs into ~/.repowiki/repos/, then scans locally |
| Scanner | src/repowiki/core/scanner.py |
Walks the tree with ignore rules and filters, and applies a priority-based 1000-file cap |
| Analyzer | src/repowiki/core/analyzer.py |
Overview, per-module, architecture and reading-guide LLM passes, cached |
| Dependency graph | src/repowiki/core/graph.py |
Regex import extraction, power-iteration PageRank, cycles, Mermaid module graph |
| Wiki builder | src/repowiki/core/wiki_builder.py |
Renders the fixed page set, sidebar and cross-links |
| Retrieval | src/repowiki/core/rag.py |
SimpleRAG chunk index and ModuleIndex card index, persisted as JSON |
| LLM client | src/repowiki/llm/ |
litellm wrapper with retries; prompt templates and extract_json |
| Server | src/repowiki/server/ |
FastAPI app, in-memory project store, SSE progress and chat |
| Exporters | src/repowiki/export/ |
Incremental Markdown, JSON, single-file HTML, GitHub Pages loader |
How a request flows
Take repowiki scan ./myrepo with an API key configured:
- Config.
scanloadsConfigfrom~/.repowiki/config.json, then appliesREPOWIKI_*env vars and CLI flags. Without an API key it prints the scan summary and stops (cli.py, config.py). - Scan.
ingest_localcallsscan_directory. It skips vendored directories, ignored paths, sensitive names, binaries, minified files and anything over 200 KB. If more thanmax_filesfiles survive, configs and entry points are kept first, then code, then everything else (scanner.py). - Overview pass.
Analyzer.analyzesends the file tree plus every config and entry-point file to the overview prompt (analyzer.py). - Module passes.
_group_into_modulesbuckets files by top-level directory._analyze_modulesruns one prompt per bucket under a semaphore (default 5) and sorts the results by size. Each prompt contains every file in the bucket, with each file cut to 4 KB or replaced by an AST skeleton for large Python files (analyzer.py, skeleton.py). - Architecture and reading guide. The architecture prompt returns Mermaid component and sequence diagrams. The reading-guide prompt gets the top 20 files by PageRank plus the module summaries (analyzer.py).
- Parse. Every pass asks for raw JSON and parses it with
extract_json, which strips fences and falls back to the outermost braces (prompts.py). A module that fails to parse gets a placeholder page and is reported as degraded. There is no re-ask. - Build.
WikiBuilder.buildassembles the fixed page list and links symbol names and file paths to the module page that owns them (wiki_builder.py). - Export.
export_markdownwrites one.mdper page plus_sidebar.mdandREADME.md. It skips pages whose input fingerprint and rendered hash match.repowiki-state.json(markdown.py).
Key components
Dependency graph and PageRank
DependencyGraph.build_from_project adds one node per file and an edge for every regex-matched import that resolves to a scanned path (graph.py). PageRank is a hand-written power iteration, so scipy is not needed (graph.py). Resolution is solid for Python and relative JS/TS imports. It is weak elsewhere: a Go import resolves only if <last two path parts>.go happens to exist, and Rust and Java rely on fixed src/ layouts (graph.py). On a Go repo, expect a mostly flat ranking.
Cache and incremental runs
Every pass result goes to ~/.repowiki/cache.db under a key of model, language, pass and content hash, with a one-year TTL (cache.py). The granularity is the module, not the file. Editing one file re-runs that file’s whole module prompt. Adding or removing any file changes the tree hash and re-runs the overview, architecture and reading-guide passes. A prompt change is only picked up after repowiki cache-clear.
Retrieval and chat
SimpleRAG cuts files into chunks of up to 30 lines and scores them by TF-IDF cosine, with CJK bigram tokens. There are no embeddings and no vector store. The index is saved under ~/.repowiki/rag/ and rebuilt per changed file (rag.py, L271-L296). ModuleIndex is meant to raise files whose LLM-written module card matches a paraphrased question (rag.py). At this commit, though, the web chat builds it from wiki.modules, and the stored Wiki dataclass has no modules field. The boost is therefore always None in the server (chat.py, wiki_builder.py). The CLI chat never passes a boost (cli.py). In practice, both chats are lexical top-5 retrieval plus up to six turns of history.
LLM client
LLMClient wraps litellm.acompletion and retries transient errors with exponential backoff. After the last retry it returns the error as a string instead of raising, and that string then fails JSON parsing upstream (client.py). One model serves every pass and the chat.
Web server
create_app mounts the scan, wiki and chat routers and serves the built SPA (app.py). Projects live in a module-level dict, so every generated wiki is lost on restart. The projects table in the SQLite cache is never written (app.py). _run_scan runs the same analyzer as the CLI, as a background task, and pushes progress lines that /status streams over SSE (scan.py).
Extending it
- Models. Any litellm
provider/modelstring works.MODEL_ALIASESandMODEL_API_BASESadd short names and OpenAI-compatible endpoints for DeepSeek, Qwen, Kimi, GLM, MiniMax and MiMo, andapi_basepoints at any other gateway (config.py). - Output language.
--langaccepts en, zh, ja or ko and only changes one sentence in each prompt. - New languages for the graph. Add regexes to
_IMPORT_PATTERNSand a branch in_resolve_import. - Page types. Add a
_build_*_pagemethod and a line inWikiBuilder.build. The layout lives in code, not configuration. - Agents.
repowiki map --format jsonandrepowiki diff --format jsonproduce prompt-ready file rankings with no LLM cost.
Running it
- Install.
pip install repowikifor the CLI, orpip install repowiki[web]for the server. Python 3.10+. On 3.10, litellm is capped below 1.98. - Key.
repowiki config set api_key ..., orREPOWIKI_API_KEY. As a fallback it readsDEEPSEEK_API_KEY,OPENAI_API_KEYorANTHROPIC_API_KEY, whichever comes first, whatever the model. The default model isdeepseek/deepseek-chat. - Generate.
repowiki scan <path|url> -f markdown|json|html. The HTML file is a single page, but it loads Mermaid from a CDN. - Serve.
repowiki serve [path]runs uvicorn on0.0.0.0(cli.py). There is no authentication, andPOST /api/scanaccepts any server-sidepath. Keep it on localhost. - Services. None beyond the LLM endpoint. The cache, RAG index and clones all live under
~/.repowiki/.
Strengths and caveats
- Strength: cheap and predictable. About three plus N LLM calls per repo (N = number of modules), aggressive caching, and a page set you can predict before running it.
- Strength: useful zero-LLM tools.
mapanddiffgive agents and reviewers a PageRank reading order, cycles and blast radius for free. - Strength: careful scanning. Priority-based capping, sensitive-file filtering and a “partial coverage” note on the overview page when files were dropped.
- Caveat: shallow understanding. Each module prompt sees at most 4 KB per file, and nothing explores the code further. Large modules are sent as one prompt with no splitting, so a big
src/<pkg>can overflow the model’s context. - Caveat: module = directory. A monorepo with one huge package gets one huge page. Ten tiny top-level folders get ten pages.
- Caveat: brittle JSON. No structured-output mode and no retry on bad JSON. Failures become placeholder pages.
- Caveat: stale clones. A cached clone under
~/.repowiki/repos/is reused forever. The CLI and server never passforce_reclone, so a URL scan does not pick up new commits (github.py). - Caveat: chat is lexical. The module-card boost is not wired up in either chat path at this commit, so paraphrased questions depend on shared words.
Sources: code at dffe687, verified Q&A.
How it answers the Open-source DeepWiki questions
Each answer was drafted by a code-reading agent at commit dffe687. Its citations were checked mechanically. Compare with the other open-source deepwiki →
How is a repository ingested and chunked?
answeredIngestion starts with ingest_local() or ingest_github() (src/repowiki/ingest/local.py:41-62, src/repowiki/ingest/github.py:70-127). For local repos, ingest_local() calls scan_directory() (src/repowiki/core/scanner.py:253-406), which walks the directory tree, filtering out:
- Well-known skip dirs (
.git,node_modules,__pycache__,dist, etc.) (scanner.py:14-20) - Binary/media/archive file extensions (
.png,.zip,.exe,.pyc,.lock, etc.) (scanner.py:22-35) - Sensitive files (
.env,.npmrc, SSH keys) (scanner.py:37-48) - Minified source files (heuristic: long lines, few non-empty) (
scanner.py:195-205) - Files larger than 200 KB, symlinks, and binary blobs (null byte check) (
scanner.py:316-343)
It also reads .gitignore and .repowikiignore via IgnoreRules (scanner.py:145-173). A hard cap of 1000 files applies; when exceeded, files are priority-sorted (configs first, then code, then docs/assets) and the lowest-priority excess is dropped (scanner.py:374-392).
For GitHub repos, ingest_github() parses github.com, gitlab.com, and bitbucket.org URLs (github.py:21-38). It runs git clone --depth 1 --single-branch with an optional GITHUB_TOKEN for private repos (github.py:98-105), checks the repo is under 500 MB (github.py:120-125), caches the clone in ~/.repowiki/repos/, and then delegates to ingest_local() (github.py:88).
Chunking for RAG (not analysis) happens in _split_into_chunks() (src/repowiki/core/rag.py:358-387): it splits file content into chunks of at most 30 lines, preferring blank-line boundaries when the chunk has at least 5 lines. Each chunk carries its file path and line range for citation.
How is retrieval (RAG) implemented?
answeredRetrieval uses pure TF-IDF with no embedding models or vector database — no external service is needed. The core class is SimpleRAG (src/repowiki/core/rag.py:52-209). At index time, every project file is split into chunks via _split_into_chunks() (rag.py:358-387), tokenized by _tokenize() (rag.py:316-328), and stored as term-frequency vectors (Counter objects) alongside an inverse-document-frequency map built by _compute_idf() (rag.py:331-340). The tokenizer handles CJK by generating character bigrams (rag.py:324-327).
Retrieval for a question: rag.retrieve(query, top_k=5) (rag.py:170-209). Each chunk gets a TF-IDF cosine similarity score against the query tokens. Results are sorted by (direct-hit flag, score), so lexical matches always rank above boost-only chunks. The top 5 chunks are returned.
A second index, ModuleIndex (rag.py:212-256), runs TF-IDF over the LLM-generated module card text (purpose, key concepts, file purposes — vocabulary not present in raw code). Its file_scores() method produces a per-file boost map. In the chat endpoint, this boost is passed into rag.retrieve() as a boost parameter (rag.py:177-178), so a paraphrased question with zero lexical overlap against code can still reach the right files through the module card match. The built-in TF-IDF index is persisted to ~/.repowiki/rag/ as JSON and loaded incrementally: only files whose content hash changed are re-tokenized (rag.py:78-121). Retrieved chunks are formatted by format_context() (rag.py:299-313) into fenced code blocks labelled with file path and line range, then injected into the chat prompt as "Relevant Code".
wiki.modules from a Wiki object that has no such field, so boost is always None, and the CLI chat never passes one; both paths are plain lexical TF-IDF.How is the wiki structure (table of contents) determined?
answeredThe wiki structure (pages and sidebar) is determined entirely by the WikiBuilder.build() method (src/repowiki/core/wiki_builder.py:42-108), which assembles pages from the structured output of the LLM analysis pipeline. It works in this order:
- Index/Overview — from
ProjectOverview(LLM-generated), always present (wiki_builder.py:54-58). - Architecture — from
ArchitectureDiagram(LLM-generated), only ifarchitecture_typeis non-empty (wiki_builder.py:60-65). - Knowledge Cards — one compact card per module, computed from
ModuleDocdata without an extra LLM call; only if any module exists (wiki_builder.py:69-72). - Module pages — one per module; modules are derived from the file tree by
_group_into_modules()in the Analyzer (src/repowiki/core/analyzer.py:147-164), which groups files by their top-level directory (or second-level if undersrc//lib/pkg/internal/app). The LLM returns oneModuleDocJSON per group with purpose, files, relationships, and concepts (analyzer.py:199-245). - Reading Guide — from
ReadingGuide(LLM-generated), only if steps exist (wiki_builder.py:88-92). - Dependencies — a Mermaid diagram of inter-module imports from
DependencyGraph, plus PageRank core files, entry points, cycles, and isolated files (wiki_builder.py:94-99). - Symbol Index — every
key_symbolsentry across all module files, grouped by kind then by module; only if any symbols were documented (wiki_builder.py:101-105).
The sidebar is a linear list with a collapsible "Modules" group containing child items. The LLM never proposes page names or hierarchy — the structure is fixed by the builder; the LLM fills the content slots.
How are individual pages generated?
answeredEach page type is generated by a dedicated method in WikiBuilder (src/repowiki/core/wiki_builder.py) that renders Pydantic model data into Markdown. The content originates from four distinct LLM passes orchestrated by Analyzer.analyze() (src/repowiki/core/analyzer.py:60-109):
Overview — prompt includes the file tree and key config/entrypoint files (prompts.py:26-56). The LLM returns JSON with name, one-liner, description, tech stack, setup instructions, and key features.
Module pages — one LLM call per module group. The prompt (prompts.py:59-97) includes the project summary and all files in that module (with full content for configs, preview/skeleton for others). The LLM returns JSON with purpose, description, per-file docs with key_symbols (name, kind, description), internal relationships, and key concepts. The builder renders this into file subsections with backtick symbol listings and relationship arrows (wiki_builder.py:237-268).
Architecture — the LLM generates Mermaid component/sequence diagrams (prompts.py:100-134). These are embedded in fenced ````mermaid blocks (wiki_builder.py:171-186). The self-contained HTML export ships the Mermaid.js CDN library and renders diagrams client-side (export/html.py:205-206`).
Module analysis runs in parallel — _analyze_modules() uses asyncio.as_completed() with a Semaphore(concurrency) (default 5) (analyzer.py:166-197). Concurrency is configurable via Config.concurrency (config.py:59).
Caching: every LLM result is stored in a SQLite cache at ~/.repowiki/cache.db, keyed by {model}:{language}:{pass_type}:{content_hash} (cache.py:27-73). On re-run, only files whose content changed (detected by content_hash()) trigger a new LLM call. The markdown export also maintains per-page fingerprints in .repowiki-state.json (state.py), so pages whose source and rendered content are unchanged are not rewritten. The --full flag forces regeneration. For oversized Python files, context_for_file() (skeleton.py:92-108) emits an AST-derived skeleton of top-level classes and functions instead of blind head-truncation.
How are model providers configured?
answeredModel providers are configured through litellm, which supports 100+ providers via a provider/model string. The Config class (src/repowiki/config.py:50-115) loads settings from three sources in order: CLI flags, environment variables (REPOWIKI_MODEL, REPOWIKI_API_KEY, REPOWIKI_API_BASE, REPOWIKI_LANG), and ~/.repowiki/config.json. If no API key is set, it falls back to provider-specific env vars: DEEPSEEK_API_KEY, OPENAI_API_KEY, or ANTHROPIC_API_KEY (config.py:85-89).
Model aliases map common names to litellm model strings (config.py:18-32): deepseek → deepseek/deepseek-chat, claude → anthropic/claude-sonnet-4-6, gpt → gpt-5.4, gemini → gemini/gemini-3.1-pro-preview, qwen → openai/qwen3.7-plus, kimi → openai/kimi-k3, mimo → openai/mimo-v2.6-flash, and others. The mimo aliases ship Xiaomi's OpenAI-compatible endpoint via MODEL_API_BASES (config.py:34-43), which also maps qwen to DashScope, kimi to Moonshot, glm to BigModel, etc. Any other OpenAI-compatible endpoint can be targeted with repowiki config set api_base https://your-gateway/v1 (config.py:98-99).
There is no per-stage model selection — a single model is used across all four analysis passes and chat. The LLMClient (src/repowiki/llm/client.py:37-144) wraps litellm's acompletion() with retry logic for transient errors (RateLimitError, APIConnectionError, Timeout, etc.) (client.py:53-68), tracks token usage and cost (client.py:97-106), and supports both complete() (non-streaming) and stream() (async generator) methods. The response_format parameter can be passed for structured output (client.py:89). litellm itself is lazily imported to avoid import cost on zero-LLM commands like repowiki map (client.py:25-34).
How is interactive Q&A / chat implemented?
answeredQ&A is available both in the terminal (repowiki chat .) and via a FastAPI streaming endpoint (/api/project/{id}/chat). Both paths follow the same pattern: ingest the repo, build/load the TF-IDF index, retrieve relevant chunks, and prompt an LLM with the context.
CLI chat (src/repowiki/cli.py:488-575): after indexing, it enters a REPL loop. The user's question goes through _answer_question() (cli.py:562-575), which calls rag.retrieve(question, top_k=5) and build_chat_prompt() to construct messages. The prompt (src/repowiki/llm/prompts.py:178-211) includes a system instruction to answer from the code and reference specific files/line numbers, the conversation history (capped at 6 most recent turns, prompts.py:175), and the relevant code chunks. Sources are printed as a file:start-end footer after each answer (cli.py:555-559).
Web chat (src/repowiki/server/routers/chat.py:18-94): the same retrieval pipeline runs, but adds a ModuleIndex boost for paraphrased questions (chat.py:43-49). The response is server-sent events (SSE): first a data: event with references (file path + line range + snippet), then streamed content tokens via LLMClient.stream(), then a {"done": true} event (chat.py:84-94). Conversation history is passed by the client as a list of ChatTurn objects and inserted into the prompt (chat.py:81-82).
The project does not implement a dedicated "deep-research" mode (multi-turn autonomous exploration) or expose its chat via MCP. The server does provide the full wiki content, file content, and dependency graph at REST endpoints (/api/project/{id}/wiki, /api/project/{id}/file/{path}, /api/project/{id}/graph) (server/routers/wiki.py:12-100), which could serve as building blocks for external integration.
wiki.modules), so both CLI and web chat use the same lexical top-5 retrieval.