LLMs Technical Reviews

How is a repository ingested and chunked?

Clone or local path; file filters; chunking strategy; supported hosts; large-repo limits.

Verdict

Both tools work on a full checkout, but they decide very differently what the model gets to see.

deepwiki-open does a shallow --depth=1 clone for GitHub, GitLab and Bitbucket (tokens injected into the URL) or reads a local path in place. iterate_files() keeps files whose extension is on an allow-list and drops the excluded dirs in repo.json. Every remaining file is split into 350-word chunks with 100 words of overlap, each tagged with its line range, and embedded. Files above about 82K tokens are skipped. There is no file-count cap, so cost and indexing time grow with the repo.

OpenDeepWiki imports Git (LibGit2Sharp, with retries), ZIP uploads or allow-listed local directories. It does not chunk anything. A scan plan limits only the directory tree shown to the catalog agent: depth 2 to 4 and about 500 files (350 for repos over 20,000 files), with .gitignore and hidden paths filtered. The agent can still open any file later with ReadFile or Grep.

Choose deepwiki-open if you want every file indexed up front and searchable by similarity. Choose OpenDeepWiki if your repos are large and you would rather have the model pick a few files than pay to embed all of them. OpenDeepWiki also handles ZIP uploads and keeps branches and commits for incremental runs.

Per-project answers

AsyncFuncAI/deepwiki-open

answered

A repository is ingested via the Repo class in api/repository.py (line 140). It supports GitHub, GitLab, Bitbucket, and local paths — each with its own _clone_from_* function that handles access tokens via URL injection (oauth2:token@ for GitLab, token@ for GitHub, x-bitbucket-api-token-auth for Bitbucket). Clones use git.Repo.clone_from with --depth=1 --single-branch for shallow clones. Local paths skip cloning entirely and are used in-place. The cloned repo path is ~/.deepwiki/repo/.

File filtering is centralized in api/config.py's iterate_files() (line 467). It walks the repo directory and applies: (1) extension whitelisting from code_extensions and doc_extensions config arrays, (2) exclusion mode using default file_filters.excluded_dirs (like node_modules, .venv, .git, dist) merged with request-level exclusions, or (3) inclusion mode when included_dirs/included_files are specified. Default excluded dirs are extensive (line 3-40 of repo.json).

Chunking happens in the RAG pipeline, not at ingestion. api/rag/pipeline.py's LineTrackingTextSplitter (line 174) uses adalflow's TextSplitter with split_by: word, chunk_size: 350, chunk_overlap: 100 (from embedder.json line 36-41). Each chunk gets annotated with its 1-based start_line/end_line within the original file. Files exceeding MAX_EMBEDDING_TOKENS * 10 (82K tokens) are skipped (line 126). The pipeline then embeds each chunk and saves to a local LocalDB pickle file at `/.adalflow/databases/(line 247-274). Large-repo limits are configurable viaMAX_CONCURRENT_WIKI_TASKS(defaults to half CPU cores) andWIKI_PAGE_CONCURRENCY(default 1) inapi/services/wiki/tasks.py` (lines 62-68).

AIDotNet/OpenDeepWiki

answered

Repository sources and hosts. Repositories are imported via RepositorySource.Parse() in RepositoryAnalyzer (src/OpenDeepWiki/Services/Repositories/RepositoryAnalyzer.cs:82) which supports three RepositorySourceType values: Git (remote Git URL), Archive (uploaded ZIP), and LocalDirectory (approved local paths). Git repos use LibGit2Sharp for clone/pull, with retries (default 3 attempts, 1 s delay). Local directories are either symlinked or copied under REPOSITORIES_DIRECTORY. ZIP archives are extracted via ZipFile.ExtractToDirectory. The config object RepositoryAnalyzerOptions (line 15) specifies the base clone directory (/data by default) and an allowlist AllowedLocalPathRoots for local imports.

File filters and scan plans. The RepositoryFileFilter class (src/OpenDeepWiki/Agents/Tools/RepositoryFileFilter.cs:6) parses .gitignore files recursively and applies their patterns via GitIgnorePatternToRegex (line 181). Hidden paths (.-prefixed directories) are automatically skipped, and binary files are filtered by extension. The scan plan resolver RepositoryScanPlanResolver (src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs:12) profiles the working directory — counting files/directories by depth, extension distributions, key files (READMEs, Makefiles, etc.) — and then applies rule-based heuristics (line 87–125). For repos with >20000 total files, tree depth is clamped to 2, max total files to 350; for 5000–20000, depth goes to 3; under 5000, depth 4. Max tree nodes budget is 1500, max files per directory 30, max total files 800.

Large-repo limits. The hardcoded budgets act as safety valves: MaxTreeNodeBudget = 1500, MaxFilesPerDirectoryBudget = 30, MaxTotalFileBudget = 800 (line 15–19). These are used in Clamp() calculations in DecideByRules() (line 87). The plan can also be set to Manual mode where overrides from the repository entity are used, or Auto where it is computed and persisted.

Chunking strategy. This project does NOT chunk source code into fixed-size blocks for embedding. Instead the entire workspace (or a bounded directory/file listing) is provided to an AI agent via tool calls. The file listing depth, max nodes, and max files per directory are the chunking controls — they limit how much of the tree is presented to the LLM in a single catalog/content generation call.

Editor's note. Correction: the effective file-tree cap is min(MaxTotalTreeFiles, 500), dropping to 350 above 20,000 files. The 800 and 1,500 values are only outer clamps. During generation the tool budget caps each read at 240 lines, each listing at 20 entries and each grep at 12 results.

How is retrieval (RAG) implemented? →