sopaco/deepwiki-rs
Rust CLI (Litho) that runs a fixed chain of LLM agents over a local repo and writes C4-style Markdown docs with Mermaid diagrams.
Overview
deepwiki-rs, which calls itself Litho, is a single Rust binary. You point it at a local checkout and it writes a folder of Markdown architecture documents. The output has a fixed shape: an overview, an architecture page, a workflow page, a boundary-interfaces page, an optional database page, and one “deep exploration” page per domain module that the LLM finds. The documents follow the C4 levels (context, container, component) and contain Mermaid diagrams.
The tool is a batch generator. It has no server, no web UI, no vector store and no chat. Every stage is a call to a “StepForwardAgent” that packs earlier results into one prompt and gets back either typed JSON or Markdown. The results live in an in-process key-value Memory for the length of one run. The only state that survives between runs is an on-disk LLM response cache keyed by an MD5 hash of the full prompt.
It is aimed at teams that want a one-command architecture write-up of a repo in one of eight output languages, and that can run a hosted or local LLM. It is not aimed at people who want to ask questions about a codebase.
Architecture
flowchart TD
CLI["cli.rs: Args -> Config"] --> L["workflow::launch"]
L --> KS["KnowledgeSyncer (optional local docs)"]
L --> P["PreProcessAgent"]
P --> SE["StructureExtractor (walk + git ls-files)"]
P --> DS["DirectorySummarizer (LLM per directory)"]
P --> RA["RelationshipsAnalyze (LLM)"]
P --> M[("Memory (in-process)")]
L --> R["ResearchOrchestrator: 6-7 research agents"]
R --> M
L --> C["DocumentationComposer: 6 editors"]
C --> M
R --> LLM["LLMClient (rig providers)"]
C --> LLM
DS --> LLM
LLM --> CACHE[("CacheManager: .litho/cache")]
LLM --> TOOLS["file_explorer / file_reader tools"]
L --> O["DiskOutlet -> output dir"]
O --> MF["mermaid-fixer (external binary)"]
| Component | Path | Role |
|---|---|---|
| CLI and config | src/cli.rs, src/config.rs |
Parse flags, load litho.toml, apply overrides and defaults |
| Pipeline driver | src/generator/workflow.rs |
launch: preprocess, research, compose, save, summary |
| Preprocessing | src/generator/preprocess/ |
File walk, README extraction, per-directory LLM dossiers, relationship analysis |
| Language processors | src/generator/preprocess/extractors/language_processors/ |
Regex extractors for 12 languages that pre-list interfaces and imports for the prompt |
| Research agents | src/generator/research/agents/ |
System context, domain modules, architecture, workflow, key modules, boundary, database |
| Editors | src/generator/compose/agents/ |
Turn research results into the final Markdown pages |
| Agent framework | src/generator/step_forward_agent.rs, agent_executor.rs |
Data-source declaration, prompt assembly, compression, cached LLM calls |
| LLM client | src/llm/client/ |
rig provider clients, model choice, retries, tool-using agents |
| Agent tools | src/llm/tools/ |
file_explorer, file_reader, time |
| Output | src/generator/outlet/ |
DocTree, DiskOutlet, SummaryOutlet, MermaidFixer |
| External knowledge | src/integrations/ |
Optional ingestion of local PDFs, Markdown, SQL and similar files by category |
How a request flows
Take deepwiki-rs -p ./my-repo -o ./litho.docs:
- Configure.
Args::to_configloadslitho.tomlfrom the working directory or--config, then lets CLI flags override provider, models, base URL, key, language and stage-skip flags (cli.rs). - Startup checks.
launchoptionally wipes the cache (--force-regenerate). It then refuses to run if themermaid-fixerbinary is not onPATH, builds theLLMClient,CacheManagerand an emptyMemory, and syncs external knowledge if configured (workflow.rs). - Preprocess.
PreProcessAgent::executereads the README, walks the tree withStructureExtractor(honouringgit ls-filesby default), asks the LLM for aDirectoryDossierper directory, and then runsRelationshipsAnalyze. All four results are stored inMemoryunder thePREPROCESSscope (preprocess/mod.rs). - Research.
ResearchOrchestratorruns the agents in a fixed order:SystemContextResearcher(C1),DomainModulesDetector,ArchitectureResearcherandWorkflowResearcher(C2),KeyModulesInsight(C3-C4),BoundaryAnalyzer, andDatabaseOverviewAnalyzeronly when a directory or file looks database-related (orchestrator.rs). - Per-agent call. Each agent’s
executechecks that its required sources exist inMemory, builds a prompt from them, appends a target-language instruction, and callsextract(typed JSON),prompt(text) orprompt_with_tools(tool-using agent) depending onLLMCallMode. The result is written back toMemory(step_forward_agent.rs). - Compose.
DocumentationComposer::executeruns the overview, architecture, workflow, key-modules, boundary and (optional) database editors in sequence (compose/mod.rs).KeyModulesInsightEditorfans out one editor per domain report, bounded bymax_parallels, and registers each page underdeep_exploration/<domain>.md(key_modules_insight_editor.rs). - Write.
DiskOutlet::savedeletes the output directory, writes everyDocTreeentry it finds inMemory, and then shells out tomermaid-fixerover the folder (outlet/mod.rs).SummaryOutletadds a run report.
Key components
Directory dossiers
Preprocessing is directory-first. generate_directory_dossiers reads each directory’s files, sorts them by name, and splits them into batches of at most 256 KB of raw content (preprocess/mod.rs). The 256 KB limit does not mean the model sees 256 KB. build_summary_prompt sends only the first 500 characters of each file. Next to them it sends a list of up to 20 interfaces and 15 imports that the regex-based language processors extracted, plus line, function and class counts (directory_summary.rs). The model returns a summary, an importance score, key files and per-file insights. If a call fails, the directory gets an empty fallback dossier and the run continues.
The 12 language processors (Rust, JS, TS, React, Vue, Svelte, PHP, Kotlin, Python, Java, C#, Swift) are line-regex extractors. There is no tree-sitter or compiler-grade parsing (language_processors/mod.rs).
StepForwardAgent
Agents declare required_sources and optional_sources (memory keys, other agents’ results, or external-knowledge categories). GeneratorPromptBuilder formats each source, runs PromptCompressor on it when it is too large, then compresses the assembled prompt once more (step_forward_agent.rs). The default formatter keeps 25 code insights, omits source code, and lists only directories when there are more than 100 files (L111-L123). Editors put __CURRENT_UTC_TIME__ placeholders in the prompt, not the real time, and substitute the time after generation. That keeps the prompt stable, so the cache can hit.
LLM client and model choice
ProviderClient::new maps eight providers to rig clients: OpenAI-compatible, Moonshot, DeepSeek, Mistral, OpenRouter, Anthropic, Gemini and Ollama (providers.rs). Model choice is a size rule: if the system and user prompts together are at most 32 KB, the run uses model_efficient with model_powerful as fallback; otherwise it uses model_powerful only (utils.rs). On a failed extraction, the fallback call repeats the prompt with the previous error appended (client/mod.rs). For the OpenAI variant, a response-parsing error falls back to a raw /chat/completions POST (providers.rs).
Tool-using agents
Three kinds of stage run as tool-using agents: ArchitectureResearcher, WorkflowEditor, and the per-domain key-module editors. AgentBuilder attaches file_explorer and file_reader and adds “Do not fabricate non-existent code” to the system prompt (agent_builder.rs). rig runs the tool loop with max_turns (default 100). file_reader refuses paths outside the project root and returns at most 200 lines unless asked for a range (file_reader.rs). ReActExecutor always reports success without a max-depth flag (react_executor.rs), so the SummaryReasoner fallback in prompt_with_react never runs at this SHA.
Cache
CacheManager stores each reply as JSON under <cache_dir>/<scope>/<md5(prompt)>.json with a timestamp and token estimate. Entries expire after expire_hours, which defaults to 8760 (one year) (cache/mod.rs). Re-runs on an unchanged repo are almost free. Any change to an upstream summary changes every downstream prompt, so those stages miss the cache.
Extending it
- New provider. Add an
LLMProvidervariant and the matching arms inProviderClient/ProviderAgent. Thematchblocks repeat for every provider, so this is a mechanical but wide change. - New page. Implement
StepForwardAgentfor a research agent and an editor, add both to the orchestrator and composer, and give the editor aDocTreeslot. The page list is code, not config. - New language. Implement
LanguageProcessor(dependencies and interfaces) and register it inLanguageProcessorManager::new. - External knowledge.
[knowledge.local_docs]inlitho.tomllets you feed PDFs, Markdown, SQL, YAML or JSON by category to chosen agents.sync-knowledgechunks and caches them.
Running it
- Install.
cargo install deepwiki-rs, pluscargo install mermaid-fixer. The run aborts at startup without mermaid-fixer (workflow.rs). - Configure. Out of the box the provider is “openai” pointed at an OpenAI-compatible ModelScope endpoint with Qwen models,
max_tokens131072,max_parallels3, and the key fromLITHO_LLM_API_KEY(config.rs). Override with--llm-provider,--model-efficient,--model-powerful,--llm-api-base-urland--llm-api-key. Ollama defaults tolocalhost:11434. - Scope. By default only git-tracked files are read.
*.md,*.txt, lock files and.envare excluded, and test files are skipped. - Output.
./litho.docsby default, with file names localised for the target language (zh, en, ja, ko, de, fr, ru, vi).
Strengths and caveats
- Strength: predictable structure. The page list and agent order are fixed in code, so two runs on similar repos give documents with the same layout.
- Strength: wide provider reach. Eight providers through rig, a two-model size rule with fallback, and an OpenAI HTTP fallback for picky compatible endpoints.
- Strength: grounded deep pages. The tool-using editors can open real files inside the project root, and they are told not to invent code.
- Strength: cheap re-runs. The prompt-hash cache plus placeholder timestamps make repeated runs on the same tree hit the cache.
- Caveat: shallow first pass. Directory dossiers see 500 characters per file plus regex-extracted signatures. Everything later builds on these summaries, so anything missed here is missed in the final pages unless a tool-using agent looks it up.
- Caveat: hard external dependency.
mermaid-fixermust be installed, and it is called with the API key as a command-line argument (fixer.rs). Other users on the machine can see that argument in the process list. - Caveat: destructive output.
DiskOutletrunsremove_dir_allon the output path before writing. Point-oat a dedicated folder. - Caveat: no incremental model.
Memoryis per-process, so--skip-preprocessingonly works when the cache can rebuild the earlier stages. There is no diff-based update of single pages. - Caveat: crude heuristics. The database stage triggers on any directory whose name contains “db”, and backend directories are explicitly rated above frontend ones in the dossier prompt.
Sources: code at 27dbd3e, verified Q&A.
How it answers the Open-source DeepWiki questions
Each answer was drafted by a code-reading agent at commit 27dbd3e. Its citations were checked mechanically. Compare with the other open-source deepwiki →
How is a repository ingested and chunked?
answeredRepository ingestion is a multi-step process orchestrated by PreProcessAgent.execute() in src/generator/preprocess/mod.rs. First, StructureExtractor.extract_structure() (src/generator/preprocess/extractors/structure_extractor.rs) scans the project directory recursively using tokio::fs::read_dir, respecting configurable filters: excluded_dirs, excluded_files, excluded_extensions, included_extensions, max_file_size (default 512KB), max_depth (default 10), include_hidden, and include_tests. When git_tracked_only is true, it runs git ls-files and skips untracked files. Binary files are detected by extension. Each file gets a heuristic importance score based on path, extension, and size. Directories also receive an LLM-boosted score via DirectoryScorer. The structure is cached under key structure_<path>.
Second, original_document_extractor.extract() (src/generator/preprocess/extractors/original_document_extractor.rs) reads README.md from the project root, stripping markdown headers and code fences, keeping the first 500 lines as a plain-text description.
Third, generate_directory_dossiers() reads each directory's files from disk (respecting all the same exclusions), batches them at 256KB total content per LLM call, and sends them to DirectorySummarizer which produces a DirectoryDossier (purpose, importance score, summary, key files). If a single batch call fails, a fallback dossier with empty summary is used. Larger directories are split into multiple lexicographically-sorted batches and merged via summarize_batch().
Finally, RelationshipsAnalyze analyzes inter-directory relationships from the dossiers. All results are stored in Memory under MemoryScope::PREPROCESS. There is no cloning step — the tool operates on a local path. Supported hosts are all filesystem directories. Large repos are handled via the 256KB batch limit, importance-based file filtering, and max_file_size truncation.
How is retrieval (RAG) implemented?
not applicableDeepWiki-rs does not implement RAG or vector retrieval. There is no embedding model, vector store, or similarity search. Instead, it uses a Memory-based data access pattern where all preprocessing and research results are stored in an in-memory Memory store (key-value with scopes) and retrieved by key lookup. Agents declare their required data sources via AgentDataConfig — e.g. DataSource::MemoryData, DataSource::ResearchResult, or DataSource::ExternalKnowledgeByCategory — and the StepForwardAgent framework assembles them into a prompt via GeneratorPromptBuilder.build_standard_user_prompt(). This is not retrieval-augmented generation in the embedding sense: the system gathers all relevant context (project structure, code insights, relationship analysis, research reports, external knowledge) and packs it directly into the LLM prompt as formatted text. Large content is compressed via PromptCompressor or emergency-truncated.
How is the wiki structure (table of contents) determined?
answeredThe document structure is hardcoded, not LLM-proposed. The DocTree in src/generator/outlet/mod.rs defines five fixed document slots matching the six compose agents: Overview, Architecture, Workflow, Boundary, and Database (plus dynamic Key Modules Insight entries). The DocTree::new() constructor creates entries mapping each AgentType to a localized filename via target_language.get_doc_filename(). The DocumentationComposer.execute() method (src/generator/compose/mod.rs) runs six editors in a fixed sequence: OverviewEditor, ArchitectureEditor, WorkflowEditor, KeyModulesInsightEditor, BoundaryEditor, and optionally DatabaseEditor.
The KeyModulesInsightEditor is the only dynamic case: it reads the KeyModulesInsight research output (a list of KeyModuleReport objects, one per domain), spawns one KeyModuleInsightEditor per domain, runs them concurrently (bounded by max_parallels), and inserts each result into the DocTree under deep_exploration/<domain_name>.md.
Research agents (system context, domain modules, architecture, workflow, boundary, database) also produce structured output with agent-defined schemas like DomainModulesReport, SystemContextReport, BoundaryAnalysisReport, etc. These feed into the document editors but do not control the TOC structure.
The output format is Markdown with Mermaid diagrams (system context diagrams, flowcharts, sequence diagrams). Editors emit instructions in their system prompts about ASCII-only node IDs and standard diagram headers.
How are individual pages generated?
answeredEach page is generated by a dedicated editor agent implementing the StepForwardAgent trait. The six editors (OverviewEditor, ArchitectureEditor, WorkflowEditor, KeyModulesInsightEditor, BoundaryEditor, DatabaseEditor) each define their own PromptTemplate with system prompt, opening instruction, closing instruction, and LLMCallMode. Most editors use LLMCallMode::Prompt (plain text generation) — only KeyModuleInsightEditor uses PromptWithTools (ReAct mode with file_explorer and file_reader tools). Per the agent prompt guidelines in overview_editor.rs, editors are instructed to include Mermaid diagrams with restrictions: ASCII-only node IDs, no zero-width characters, standard diagram headers. After generation, MermaidFixer (src/generator/outlet/fixer.rs) runs mermaid-fixer as an external program on the output directory to auto-fix any syntax issues.
Parallelism is implemented in KeyModulesInsightEditor.execute(): it creates one KeyModuleInsightEditor per domain report and runs them concurrently using do_parallel_with_limit() from src/utils/threads.rs (a semaphore-bounded concurrent executor). The concurrency limit is config.llm.max_parallels (default 3).
Caching is performed at the LLM call level in agent_executor.rs (extract(), prompt(), prompt_with_tools() functions). Each function computes a cache key from the system+user prompt, checks the CacheManager, and stores results with token usage estimates. Cache scope is the agent type key, so re-running with identical prompts hits the cache. --force-regenerate clears the cache before starting. Output is written by DiskOutlet.save() which iterates the DocTree, retrieves each document from Memory by its scoped key, and writes it to output_path. Source file citations come from the research data fed into each editor's prompt.
WorkflowEditor also runs in PromptWithTools mode alongside the per-domain KeyModuleInsightEditor; the other editors (overview, architecture, boundary, database) use plain Prompt.How are model providers configured?
answeredModel providers are configured through the LLMConfig struct in src/config.rs. The LLMProvider enum supports 8 providers: OpenAI, Moonshot, DeepSeek, Mistral, OpenRouter, Anthropic, Gemini, and Ollama. Each is mapped to a rig provider client in src/llm/client/providers.rs via the ProviderClient enum. The ProviderClient::new() factory method creates the appropriate rig client, using builder patterns for custom base URLs. The only hardcoded URL check is for Anthropic: custom URLs are only used if they contain "anthropic" — otherwise it defaults to https://api.anthropic.com to avoid accidentally routing through a non-Anthropic endpoint when a global default URL is set.
Per-stage model selection is governed by evaluate_befitting_model() in src/llm/client/utils.rs: if combined system+user prompt is ≤32KB, it uses model_efficient with model_powerful as fallback; beyond 32KB, it uses model_powerful with no fallback. Both models come from config (model_efficient, model_powerful).
OpenAI-compatible endpoints are supported generically: OpenAI, Moonshot, and DeepSeek clients all accept custom base_url. For the OpenAI provider specifically, ProviderAgent::prompt() has an HTTP fallback path (prompt_via_http): when the rig agent fails with an API response parsing error, it retries via a raw reqwest POST to {base_url}/chat/completions with the same system/user messages, max_tokens, and temperature.
Ollama is supported with a dedicated ollama_extractor_wrapper that wraps the rig agent for structured extraction. Local models use api_base_url defaulting to http://localhost:11434 when provider is Ollama (set in cli.rs line 193-195). All providers support configurable max_tokens, temperature (optional), retry_attempts (default 3), retry_delay_ms (default 5000), and timeout_seconds (default 300).
How is interactive Q&A / chat implemented?
not applicableDeepWiki-rs does not implement interactive Q&A, chat, or conversational features. There is no streaming endpoint, no MCP server, no conversation memory for user interaction, and no deep-research chat mode. The ReAct executor in src/llm/client/react.rs uses chat history internally for multi-turn tool-calling within a single agent call (tracking tool call→response cycles), but this is an internal mechanism for the LLM to use file_explorer/file_reader tools during generation — not exposed to users. The SummaryReasoner similarly replays the ReAct chat history to produce a final summary when the agent hits its max turn limit. These are internal implementation details, not interactive chat features. The tool is a batch documentation generator only.