OpenBMB/RepoAgent
Python CLI that writes one LLM doc per function and class, ordered by a Jedi call graph and refreshed incrementally from git changes.
Overview
RepoAgent (from OpenBMB) generates API-level documentation for Python repositories. It does not write a wiki of concept pages. It writes one LLM-authored entry for every function, method, nested function and class it finds with ast, and assembles them into one Markdown file per .py file under markdown_docs/. Its distinctive idea is dependency order: Jedi resolves who calls whom, and the scheduler documents callees before callers, so each prompt can include the finished docs of the code it calls.
The second idea is maintenance. A checkpoint folder, .project_doc_record/, stores the whole object tree with docs, code and reference lists. On each later run RepoAgent rebuilds the tree, compares it with the checkpoint, and regenerates only objects whose code changed or whose callers changed. The README pitches this as a pre-commit hook: the hook regenerates docs and git adds them.
An optional extra, chat-with-repo, indexes the generated docs in Chroma and serves a single-turn RAG demo in Gradio. The pinned commit is from December 2024 and the package is version 0.2.0.
Architecture
flowchart TD
CLI["repoagent run (click)"] --> S["SettingsManager (pydantic-settings)"]
CLI --> R["Runner"]
R --> FF["make_fake_files (git index vs worktree)"]
R --> MI["MetaInfo.init_meta_info"]
MI --> FH["FileHandler: ast walk of .py files"]
MI --> T["DocItem tree (dir/file/class/function)"]
R --> REF["parse_reference (Jedi)"]
REF --> T
R --> TM["TaskManager (topological)"]
TM --> W["worker threads x max_thread_count"]
W --> CE["ChatEngine.generate_doc (OpenAILike)"]
CE --> T
R --> CP[(".project_doc_record checkpoint")]
R --> MD["markdown_refresh -> markdown_docs/"]
MD --> GIT["git add generated docs"]
CP --> CHAT["chat-with-repo: Chroma + Gradio"]
| Component | Path | Role |
|---|---|---|
| CLI | repo_agent/main.py |
run, diff, clean, chat-with-repo commands |
| Settings | repo_agent/settings.py |
Project and chat-completion settings from flags or env |
| Runner | repo_agent/runner.py |
First generation, incremental update, Markdown output, git staging |
| Object tree | repo_agent/doc_meta_info.py |
DocItem, MetaInfo, Jedi references, topology, checkpoint merge |
| Parsing | repo_agent/file_handler.py, utils/gitignore_checker.py |
.gitignore-aware walk and AST extraction |
| Git staging trick | repo_agent/utils/meta_info_utils.py |
Swaps unstaged edits out so docs match the index |
| Scheduler | repo_agent/multi_task_dispatch.py |
Dependency-aware task queue and worker loop |
| LLM call | repo_agent/chat_engine.py, prompt.py |
Prompt assembly and one chat call per object |
| Chat demo | repo_agent/chat_with_repo/ |
Chroma index, multi-query RAG, Gradio UI |
How a request flows
Take a first repoagent run -tp ./myrepo:
- Settings.
runpasses CLI options intoSettingsManager.initialize_with_params. The API key comes only from theOPENAI_API_KEYenvironment variable (main.py, settings.py). - Build the tree.
Runner.__init__sees no checkpoint, callsmake_fake_files, thenMetaInfo.init_meta_info, and writes the first checkpoint (runner.py).generate_overall_structurewalks non-ignored.pyfiles and runsget_functions_and_classeson each (file_handler.py, L258-L305).from_project_hierarchy_jsonnests objects by line-range containment, so the smallest enclosing object becomes the parent (doc_meta_info.py). - Resolve references.
get_topologycallsparse_reference. For every object, Jediget_referencesat its name position yields referencing locations. Each location is mapped back to the enclosingDocItem, which fillsreference_whoandwho_reference_me(doc_meta_info.py, L500-L615). - Schedule.
get_task_managerrepeatedly picks an item whose children and callees are already scheduled. A parent therefore waits for its methods, and a caller waits for its callees. On a cycle, it picks the item that breaks the fewest non-special edges and logscircle-reference(doc_meta_info.py). - Generate.
first_generatestartsmax_thread_countthreads runningworker. Each thread pollsget_next_taskfor a task with no open dependencies (runner.py, multi_task_dispatch.py).generate_doc_for_a_single_itemcallsChatEngine.generate_doc, appends the reply tomd_content, marks the itemdoc_up_to_date, and re-checkpoints after every object (runner.py). - Render.
markdown_refreshdeletesmarkdown_docs/. It then writes one.mdper source file, with heading depth following nesting, and uses the latest entry in each object’smd_content(runner.py). - Later runs. With a checkpoint present,
runbuilds a fresh tree and callsload_doc_from_older_meta, which copies docs across. It setscode_changedwherecode_contentdiffers, andadd_new_referencerorreferencer_not_existwhere the set of callers changed. Only those items are dispatched, after which the docs are staged withgit add(runner.py, doc_meta_info.py).
Key components
DocItem tree and checkpoint
Every node has a type (_repo, _dir, _file, _class, _class_function, _function, _sub_function), its source and line range, a status, a list of generated docs, and both reference lists. Only objects below file level are ever sent to the LLM. need_to_generate skips files, directories, the repo node, anything already up to date, and any path that starts with an --ignore-list prefix (doc_meta_info.py). The checkpoint is a JSON dump of this tree. The optional chat demo reads the same file.
Prompt
ChatEngine.build_prompt fills one system template with the object’s code, its type and name, a request for an output example when the function returns something, and two blocks. One block holds the code and latest doc of every callee; the other holds the code and latest doc of every caller (chat_engine.py). The template requires bold field labels and forbids Markdown headings, so the Runner can nest the replies under its own headings (prompt.py). The template also has a {project_structure} slot introduced by “the related hierarchical structure of this project is as follows”. build_prompt never passes a value for it, so at this SHA the project tree does not reach the model.
Staging trick
make_fake_files reads git status. It skips untracked .py files entirely. For every modified-but-unstaged file, it renames the working copy to <name>_latest_version.py and writes the index version back under the original name (meta_info_utils.py). Docs therefore describe what is about to be committed, which suits a pre-commit hook. delete_fake_files restores the working copy afterwards (L82-L107).
LLM access
The generator uses llama-index OpenAILike with is_chat_model=True and max_retries=1. Model, temperature, timeout and base URL come from settings (chat_engine.py). Any OpenAI-compatible server works, but there is no other provider abstraction. A failed call is logged, and the object stays undocumented until the next run.
chat-with-repo
main() loads project_hierarchy.json, embeds the docs with text-embedding-3-large into a persistent Chroma collection using a semantic splitter, and opens a Gradio UI (chat_with_repo/main.py, vector_store_manager.py). For each question, RepoAssistant.respond runs a fixed sequence (rag.py):
- extract keywords and generate query variants;
- query Chroma for each variant;
- rerank the hits with
gpt-4o-minias a JSON relevance judge; - draft an answer;
- pull class and function names out of the draft and the question, and look up their source;
- write the final answer with
gpt-4o.
Both model names are hard-coded (rag.py). The pipeline keeps no history and does not stream.
Extending it
- Prompt and format. Edit
prompt.pyandChatEngine.build_prompt. Everything about doc style lives there. - Model endpoint.
--modeland--base-urlpoint the generator at any OpenAI-compatible server. Local models work only behind such a server. - Scope.
--ignore-listprefixes and.gitignorecontrol what is parsed.--languageaccepts any ISO 639 name or code for the output language. - Another language than Python. This would mean replacing
FileHandler’s AST extraction and the Jedi reference pass. Both are Python-specific throughout.
Running it
- Install.
pip install repoagent(Python 3.11+). Addrepoagent[chat-with-repo]for the chat demo. - Configure. Export
OPENAI_API_KEY. Defaults aregpt-4o-mini, temperature 0.2, a 60 s timeout, and four threads. - Run.
repoagent run -tp <repo>generates or updates docs.repoagent diffpreviews what would change,repoagent cleanremoves leftover_latest_version.pyfiles, andrepoagent chat-with-repostarts the Gradio demo. - Requirements. The target must be a git repository with at least one commit, because the run stores
HEAD’s hash as the doc version.
Strengths and caveats
- Strength: call-graph context. Documenting callees first and feeding their finished docs into callers’ prompts is a real design idea, and it is implemented end to end with Jedi.
- Strength: incremental updates. The checkpoint diff regenerates only changed code and objects whose callers changed, which keeps repeated runs cheap.
- Strength: commit-aligned docs. The index-swap trick makes generated docs match staged code.
- Caveat: Python only, API-level only. There are no architecture pages, overviews, or diagrams, and module- and file-level docs are never generated.
- Caveat: the working tree is modified during a run. Unstaged edits are renamed aside while the run is in progress. The first-generation path returns without calling
delete_fake_files(runner.py), so after a first run with unstaged changes you needrepoagent clean. - Caveat: scaling. Jedi reference lookup runs once per object, the scheduler is quadratic in the number of objects, and the full checkpoint is rewritten after every generated doc.
- Caveat: dated and partly dead code.
process_file_changes,add_new_itemandupdate_existing_iteminrunner.pyare an older path that nothing calls. The chat demo hard-codes OpenAI model names and passes the project name “test” into its final prompt.
Sources: code at 825d988, verified Q&A.
How it answers the Open-source DeepWiki questions
Each answer was drafted by a code-reading agent at commit 825d988. Its citations were checked mechanically. Compare with the other open-source deepwiki →
How is a repository ingested and chunked?
answeredRepoAgent ingests a local filesystem path to a Python repository — there is no remote cloning. The path is specified via --target-repo-path / -tp, defaulting to path/to/your/target/repository (repo_agent/main.py:129).
File filtering works through two layers:
- A
GitignoreCheckerreads.gitignorepatterns, splitting them into folder and file patterns; onos.walkit prunes ignored directories entirely and skips non-.pyfiles (repo_agent/utils/gitignore_checker.py:99-124). - An additional
--ignore-list/-iCLI flag supplies extra file-path prefixes to skip (repo_agent/main.py:109); this is checked insideneed_to_generate()indoc_meta_info.py:96-106. - Unstaged new/modified files are handled via a "fake file" mechanism: Git's index content for modified-but-unstaged files is written into a temp
_latest_version.pywhile the working-copy is renamed, so the scanner sees the committed version plus the uncommitted diff separately (repo_agent/utils/meta_info_utils.py:13-79). Untracked.pyfiles are added to ajump_fileslist and skipped entirely.
Chunking strategy is per-function/per-class, not fixed-token windows. FileHandler.generate_overall_structure() walks every non-ignored .py file, parses it with ast.parse(), and extracts every FunctionDef, ClassDef, and AsyncFunctionDef plus their source lines and parameter lists (file_handler.py:176-215, 258-305). The result is a flat dict of file paths to object lists, later assembled into a tree by MetaInfo.from_project_hierarchy_json() based on line-number containment (doc_meta_info.py:932-998).
Supported hosts: the target repo must be a local directory with a git history (for diff detection). There is no GitHub/GitLab/remote-host abstraction. Only Python is supported — AST-based parsing means no other language can be ingested.
Large-repo limits: the --max-thread-count / -mtc flag (default 4) controls how many doc-generation tasks run concurrently (main.py:120-123). The reference-parsing step (using Jedi) walks every file and every object, which could be slow for very large repos, but there is no explicit file-size or object-count cap. The topology-based task scheduler avoids circular-reference deadlocks with a best-effort fallback (doc_meta_info.py:637-676).
--target-repo-path defaults to an empty string, not a placeholder path. For modified-but-unstaged files the working copy is renamed to <name>_latest_version.py and the index version is written under the original name, so documentation is generated from the staged code.How is retrieval (RAG) implemented?
answeredRetrieval (RAG) is implemented only in the optional chat-with-repo sub-package installed via pip install repoagent[chat-with-repo]. The main documentation-generation pipeline does not use RAG at all — it retrieves code context statically via the DocItem tree and reference relationships.
Within chat-with-repo, retrieval works as follows:
Embedding model: OpenAI's text-embedding-3-large, configured in vector_store_manager.py:52-55. It uses the same api_key and api_base as the chat completion.
Vector store: ChromaDB (persistent, stored at ./chroma_db). Documents are indexed into a ChromaVectorStore via LlamaIndex's VectorStoreIndex (vector_store_manager.py:48-113). The index is built by extracting markdown content and metadata from the project_hierarchy.json file, which is the output of the main doc-generation pipeline.
Chunking for RAG: The vector store uses SemanticSplitterNodeParser (buffer_size=1, breakpoint_percentile_threshold=95) with a SentenceSplitter fallback at chunk_size=1024 (vector_store_manager.py:60-63).
Top-k: default 5 (rag.py:34: top_k=5). On each query, the vector store is queried, then results are reranked by an LLM (gpt-4o-mini) that scores each document's relevance on a 0–100 scale via JSON output (rag.py:44-55). The top-5 reranked documents are kept.
How retrieved code reaches the prompt: The RepoAssistant.respond() method (rag.py:84-176) orchestrates a multi-stage pipeline:
- Extracts keywords via
TextAnalysisTool.keyword()(text_analysis_tool.py:13-16) - Generates 3 query variants via
generate_queries()(rag.py:36-42) - Queries Chroma with each variant, deduplicates by text content
- Reranks by LLM relevance score
- Runs a first-pass RAG generation via
rag()(gpt-4o-mini) - Extracts named entities (class/function names) from bot and user text via
nerquery(), then looks up raw source code from the JSON (text_analysis_tool.py:42-55) - Passes both the retrieved documents AND the raw source code into a final RAG-AR (Advanced RAG) call using the strong model (gpt-4o) (
rag.py:74-82)
The final RAG-AR prompt includes related_code (source code snippets) and embedding_recall (documentation text) as separate sections (chat_with_repo/prompt.py:43-64).
How is the wiki structure (table of contents) determined?
answeredThe wiki/documentation structure is not determined by an LLM prompt or agent. Instead, it is derived deterministically from the filesystem and AST analysis:
File tree walk:
FileHandler.generate_overall_structure()walks the target repo directory tree, respecting.gitignore, and collects every.pyfile that is not ignored (file_handler.py:258-305). The result is a plain dict keyed by relative file path.AST-based object extraction: Each
.pyfile is parsed withast.parse(), and everyFunctionDef,ClassDef, andAsyncFunctionDefis extracted with start/end line numbers, parameter lists, and source code (file_handler.py:176-215).Tree construction:
MetaInfo.from_project_hierarchy_json()(doc_meta_info.py:871-1019) assembles the flat dict into a nested tree:- File paths are split on
/to build directory nodes (_dir, _file, _repo types) - Objects within each file are placed using line-number containment: if object A's lines contain object B's lines, B becomes A's child in the tree (
doc_meta_info.py:954-998) - Object types are assigned:
ClassDef→_class,FunctionDefnested in a class →_class_function,FunctionDefin a function →_sub_function, regular →_function(doc_meta_info.py:999-1014)
- File paths are split on
Reference relationships: After the tree is built,
MetaInfo.parse_reference()(doc_meta_info.py:500-615) uses Jedi (jedi.Script.get_references()) to find all bidirectional call/reference relationships between objects across files. These populatereference_who(who this object calls) andwho_reference_me(who calls this object) on eachDocItemnode.Task scheduling: A topology-based
TaskManager(multi_task_dispatch.py) sorts objects so that callees are documented before callers, usingreference_whoas dependencies (doc_meta_info.py:617-696). Circular references are handled via a best-effort fallback (doc_meta_info.py:637-676).GitBook display: The
display/book_tools/scripts copy the generated markdown docs into a GitBook-compatible folder and auto-generate aSUMMARY.md(table of contents) by recursively scanning the markdown directory structure (display/book_tools/generate_summary_from_book.py). The TOC mirrors the filesystem hierarchy of the markdown docs, which in turn mirrors the original repo structure.
There is no LLM deciding section names or page layouts — the structure is purely code-driven.
How are individual pages generated?
answeredPer-page generation is done by ChatEngine.generate_doc() in chat_engine.py:116-132. Each DocItem (representing a single function, class method, or sub-function) gets its own LLM call. Prompts are built by ChatEngine.build_prompt() (chat_engine.py:27-114), which assembles:
- The object's source code
- Its file path within the project
- A hierarchical project structure with the current object marked (
✳️) - Code and documentation of objects it calls (
reference_who) - Code and documentation of objects that call it (
who_reference_me) - Whether the function has a return value (to conditionally request an "Output Example" section)
- A
languageparameter to control output language
The prompt follows a rigid template (prompt.py:5-32) requiring bold section headers (**function_name**, parameters_or_attribute, Code Description, Note) and forbidding the model from outputting markdown headings (so the output is a flat block that gets rendered at the correct heading level later).
Concurrency: Generation runs with max_thread_count threads (default 4). Tasks are dispatched by the TaskManager (multi_task_dispatch.py:103-125), which respects dependency ordering (callees before callers for reference content) while allowing independent objects to run in parallel. The worker loop grabs the next available task, calls generate_doc_for_a_single_item(), which calls ChatEngine.generate_doc(), then marks the task complete (runner.py:76-100, runner.py:119-155).
Caching and regeneration: The MetaInfo class persists to a checkpoint directory (.project_doc_record/) containing project_hierarchy.json and meta-info.json (doc_meta_info.py:393-439). On subsequent runs, Runner.__init__() loads this checkpoint and MetaInfo.load_doc_from_older_meta() (doc_meta_info.py:716-804) compares old vs new code content. Objects whose code_content changed get status code_changed; objects whose set of referencers changed get add_new_referencer or referencer_not_exist; unchanged objects remain doc_up_to_date and are skipped. Only objects requiring generation are dispatched as tasks.
Markdown rendering: After generation, Runner.markdown_refresh() (runner.py:157-231) iterates all file-node DocItem objects, serializes each child (class or function) as a markdown section with heading levels determined by nesting depth, and writes one .md file per original .py file. There are no Mermaid diagrams or any visual elements added — the output is pure text and markdown headings.
Diagrams: No diagram generation (Mermaid or otherwise) is implemented anywhere in the codebase.
{project_structure} placeholder, but build_prompt never fills it, so no project hierarchy is sent; the prompt contains the object's code plus the code and latest docs of its callees and callers.How are model providers configured?
answeredRepoAgent supports OpenAI-compatible endpoints via a single model provider abstraction. There is no multi-provider abstraction layer (no LangChain, no LiteLLM, no custom provider registry).
Configuration: Providers are configured through two settings classes in settings.py:
ChatCompletionSettings(settings.py:58-69):model(defaultgpt-4o-mini),temperature(default 0.2),request_timeout(default 60),openai_base_url(defaulthttps://api.openai.com/v1),openai_api_key(loaded from env, SecretStr field).ProjectSettings(settings.py:27-56):language,max_thread_count,ignore_list, etc.
These are parsed from environment variables via pydantic-settings by default, or set explicitly via CLI flags passed to SettingsManager.initialize_with_params() (settings.py:87-123).
LLM instantiation: The main doc-generation pipeline uses llama-index's OpenAILike class (chat_engine.py:17-25), which accepts any OpenAI-compatible base URL. This means any service that exposes an OpenAI-compatible API (Azure OpenAI, Ollama with an OpenAI proxy, vLLM, etc.) can be used by setting --base-url. The API key is passed through to this class.
Per-stage model selection: There are two distinct model tiers hardcoded in the chat-with-repo feature:
weak_model=gpt-4o-mini: used for query generation, keyword extraction, initial RAG response, and relevance reranking (rag.py:22-25)strong_model=gpt-4o: used only for the final RAG-AR (advanced retrieval-augmented generation) call (rag.py:26-31)
In the main doc-generation pipeline, only one model is used — whatever is configured in ChatCompletionSettings (chat_engine.py:17-25). There is no per-object-type or per-stage model selection.
Local models: The README claims "Local model support like Llama, chatGLM, Qwen, GLM4" (README.md:214), but the code contains no local-model abstraction. Since OpenAILike accepts any base URL, locally hosted models behind an OpenAI-compatible server (e.g., Ollama's OpenAI proxy, LocalAI, vLLM) would work transparently, but there is no built-in support for HuggingFace transformers, llama.cpp Python bindings, or any local inference engine. This is a README claim backed only by the generic OpenAI-compatible endpoint design.
Key management: The OpenAI API key is read from the openai_api_key environment variable (secret field, excluded from serialization) (settings.py:63).
How is interactive Q&A / chat implemented?
answeredInteractive Q&A is implemented through the chat-with-repo sub-package, activated by pip install repoagent[chat-with-repo] and launched via repoagent chat-with-repo (main.py:224-239).
Startup flow (chat_with_repo/main.py:9-43): The RepoAssistant is initialized with API credentials and the path to project_hierarchy.json. It then:
- Extracts all markdown content and metadata from the JSON via
JsonFileProcessor.extract_data()(json_handler.py:20-48) - Creates a ChromaDB vector store from these documents using
VectorStoreManager.create_vector_store()(vector_store_manager.py:31-116) - Launches a Gradio web interface (
GradioInterface,gradio_interface.py:7-195)
Chat flow (rag.py:84-176): The respond() method receives user message and optional instruction. The pipeline:
- Formats a prompt and extracts 1–3 code keywords via
TextAnalysisTool.keyword()(using gpt-4o-mini) - Generates 3 query variants via
generate_queries()(gpt-4o-mini) - Queries the Chroma vector store (OpenAI
text-embedding-3-largeembeddings, top-5) (vector_store_manager.py:117-132) - Reranks results by LLM-gated relevance scores and passes to a first-pass RAG response (gpt-4o-mini) (
rag.py:58-63) - Extracts class/function names from both user and bot text via
nerquery()(text_analysis_tool.py:42-55), then looks up actual source code from the JSON database - Runs a final RAG-AR call using gpt-4o (
rag.py:74-82), feeding it both the retrieved documentation chunks AND the raw source code, producing the final answer
Deep-research mode: No deep-research / multi-step reasoning mode exists in the codebase. The pipeline is a fixed single-pass sequence.
Streaming: The Gradio interface does NOT use streaming. The submit button produces the full response text at once (gradio_interface.py:180-192).
Conversation memory: There is no conversation memory or history tracking. Each invocation of respond() is stateless — the prompt and instruction textbox are sent fresh, and the output is rendered as markdown. The Gradio UI has a "record" button that is never wired to any backend handler.
MCP exposure: No MCP (Model Context Protocol) server or tool is implemented. The chat-with-repo feature is a standalone Gradio app, not an MCP service.
In summary, chat-with-repo is a demo-quality prototype: a single-turn RAG pipeline over pre-generated documentation, with no conversation history, no streaming, no multi-turn refinement, and no agentic capabilities beyond the fixed retrieval + generation sequence.