# How is a repository ingested and chunked?

> Open-source DeepWiki — a good answer covers: Clone or local path; file filters; chunking strategy; supported hosts; large-repo limits.

Canonical page: https://llms-technical-reviews.com/open-source-deepwiki/q/ingestion/

## Verdict

Both tools work on a full checkout, but they decide very differently what the model gets to see.

[deepwiki-open](/p/deepwiki-open/) does a shallow `--depth=1` clone for GitHub, GitLab and Bitbucket (tokens injected into the URL) or reads a local path in place. `iterate_files()` keeps files whose extension is on an allow-list and drops the excluded dirs in `repo.json`. Every remaining file is split into 350-word chunks with 100 words of overlap, each tagged with its line range, and embedded. Files above about 82K tokens are skipped. There is no file-count cap, so cost and indexing time grow with the repo.

[OpenDeepWiki](/p/opendeepwiki/) imports Git (LibGit2Sharp, with retries), ZIP uploads or allow-listed local directories. It does not chunk anything. A scan plan limits only the directory tree shown to the catalog agent: depth 2 to 4 and about 500 files (350 for repos over 20,000 files), with `.gitignore` and hidden paths filtered. The agent can still open any file later with `ReadFile` or `Grep`.

Choose deepwiki-open if you want every file indexed up front and searchable by similarity. Choose OpenDeepWiki if your repos are large and you would rather have the model pick a few files than pay to embed all of them. OpenDeepWiki also handles ZIP uploads and keeps branches and commits for incremental runs.

## Per-project answers

### AsyncFuncAI/deepwiki-open (answered)

A repository is ingested via the `Repo` class in `api/repository.py` (line 140). It supports GitHub, GitLab, Bitbucket, and local paths — each with its own `_clone_from_*` function that handles access tokens via URL injection (`oauth2:token@` for GitLab, `token@` for GitHub, `x-bitbucket-api-token-auth` for Bitbucket). Clones use `git.Repo.clone_from` with `--depth=1 --single-branch` for shallow clones. Local paths skip cloning entirely and are used in-place. The cloned repo path is `~/.deepwiki/repo/`.

File filtering is centralized in `api/config.py`'s `iterate_files()` (line 467). It walks the repo directory and applies: (1) extension whitelisting from `code_extensions` and `doc_extensions` config arrays, (2) exclusion mode using default `file_filters.excluded_dirs` (like `node_modules`, `.venv`, `.git`, `dist`) merged with request-level exclusions, or (3) inclusion mode when `included_dirs`/`included_files` are specified. Default excluded dirs are extensive (line 3-40 of `repo.json`).

Chunking happens in the RAG pipeline, not at ingestion. `api/rag/pipeline.py`'s `LineTrackingTextSplitter` (line 174) uses adalflow's `TextSplitter` with `split_by: word`, `chunk_size: 350`, `chunk_overlap: 100` (from `embedder.json` line 36-41). Each chunk gets annotated with its 1-based `start_line`/`end_line` within the original file. Files exceeding `MAX_EMBEDDING_TOKENS * 10` (~82K tokens) are skipped (line 126). The pipeline then embeds each chunk and saves to a local `LocalDB` pickle file at `~/.adalflow/databases/` (line 247-274). Large-repo limits are configurable via `MAX_CONCURRENT_WIKI_TASKS` (defaults to half CPU cores) and `WIKI_PAGE_CONCURRENCY` (default 1) in `api/services/wiki/tasks.py` (lines 62-68).


Citations: [api/config.py:467-516](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config.py#L467-L516) · [api/config/repo.json:3-150](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config/repo.json#L3-L150) · [api/rag/pipeline.py:174-275](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/rag/pipeline.py#L174-L275) · [api/services/wiki/tasks.py:62-68](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/services/wiki/tasks.py#L62-L68) · [api/config/embedder.json:33-41](https://github.com/AsyncFuncAI/deepwiki-open/blob/d92819a9c9f3b99416e3580ff235fc9d3adf8b89/api/config/embedder.json#L33-L41)

### AIDotNet/OpenDeepWiki (answered)

**Repository sources and hosts.** Repositories are imported via `RepositorySource.Parse()` in `RepositoryAnalyzer` (src/OpenDeepWiki/Services/Repositories/RepositoryAnalyzer.cs:82) which supports three `RepositorySourceType` values: `Git` (remote Git URL), `Archive` (uploaded ZIP), and `LocalDirectory` (approved local paths). Git repos use LibGit2Sharp for clone/pull, with retries (default 3 attempts, 1 s delay). Local directories are either symlinked or copied under `REPOSITORIES_DIRECTORY`. ZIP archives are extracted via `ZipFile.ExtractToDirectory`. The config object `RepositoryAnalyzerOptions` (line 15) specifies the base clone directory (`/data` by default) and an allowlist `AllowedLocalPathRoots` for local imports.

**File filters and scan plans.** The `RepositoryFileFilter` class (src/OpenDeepWiki/Agents/Tools/RepositoryFileFilter.cs:6) parses `.gitignore` files recursively and applies their patterns via `GitIgnorePatternToRegex` (line 181). Hidden paths (`.`-prefixed directories) are automatically skipped, and binary files are filtered by extension. The scan plan resolver `RepositoryScanPlanResolver` (src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs:12) profiles the working directory — counting files/directories by depth, extension distributions, key files (READMEs, Makefiles, etc.) — and then applies rule-based heuristics (line 87–125). For repos with >20000 total files, tree depth is clamped to 2, max total files to 350; for 5000–20000, depth goes to 3; under 5000, depth 4. Max tree nodes budget is 1500, max files per directory 30, max total files 800.

**Large-repo limits.** The hardcoded budgets act as safety valves: `MaxTreeNodeBudget = 1500`, `MaxFilesPerDirectoryBudget = 30`, `MaxTotalFileBudget = 800` (line 15–19). These are used in `Clamp()` calculations in `DecideByRules()` (line 87). The plan can also be set to `Manual` mode where overrides from the repository entity are used, or `Auto` where it is computed and persisted.

**Chunking strategy.** This project does NOT chunk source code into fixed-size blocks for embedding. Instead the entire workspace (or a bounded directory/file listing) is provided to an AI agent via tool calls. The file listing depth, max nodes, and max files per directory are the chunking controls — they limit how much of the tree is presented to the LLM in a single catalog/content generation call.

> **Editor's note.** Correction: the effective file-tree cap is `min(MaxTotalTreeFiles, 500)`, dropping to 350 above 20,000 files. The 800 and 1,500 values are only outer clamps. During generation the tool budget caps each read at 240 lines, each listing at 20 entries and each grep at 12 results.

Citations: [src/OpenDeepWiki/Services/Repositories/RepositoryAnalyzer.cs:82-240](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryAnalyzer.cs#L82-L240) · [src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs:12-125](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryScanPlanResolver.cs#L12-L125) · [src/OpenDeepWiki/Agents/Tools/RepositoryFileFilter.cs:6-258](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Agents/Tools/RepositoryFileFilter.cs#L6-L258) · [src/OpenDeepWiki/Services/Repositories/RepositoryScanPlan.cs:1-48](https://github.com/AIDotNet/OpenDeepWiki/blob/d33113c88c56458c202fa6904ec114387319f1f5/src/OpenDeepWiki/Services/Repositories/RepositoryScanPlan.cs#L1-L48)
