LLMs Technical Reviews

aibox22/readmeX

Python CLI that has an OpenAI-compatible LLM fill a README template or a fixed eight-page MkDocs site, with optional code RAG over Python.

GitHub ↗★ 630PythonMITcommit b83df61 · 2025-08-01

Overview

readmeX is a command-line tool with two modes. The default mode writes a README.md for a local project. It summarises source files one by one with an LLM, asks for a description, entry file, key features and extra notes, optionally generates a logo with an image model, and has the LLM fill a bundled README template. --website mode builds a MkDocs Material site with a fixed set of pages: home (the generated README), installation, usage, an API reference, examples, architecture (with a draw.io diagram), contributing and changelog. --serve and --deploy run mkdocs serve or deploy to GitHub Pages.

It describes itself as a Chinese-first open-source DeepWiki, and the website mode matches that. Page prompts, RAG queries, navigation labels and the MkDocs theme language are hard-coded in Chinese. The README mode writes English unless Chinese is chosen. There is no chat or Q&A interface. Each run is a one-shot batch job.

The code is two large modules, core.py (README, 1.5k lines) and website_core.py (site, 3.2k lines), plus a 950-line code_rag.py.

Architecture

flowchart TD
  CLI["readmex CLI (argparse)"] --> RM["readmex.generate (README mode)"]
  CLI --> WG["WebsiteGenerator.generate_website"]
  RM --> LA["LanguageAnalyzer (primary language)"]
  RM --> SD["_get_script_descriptions (1 LLM call per file)"]
  RM --> GEN["description / entry / features / extra (LLM)"]
  RM --> LOGO["generate_logo (LLM + T2I)"]
  RM --> TPL["BLANK_README.md template -> LLM -> README.md"]
  WG --> AN["_analyze_project (AST, deps, git)"]
  WG --> RAG["CodeRAG: AST blocks + embeddings + FAISS"]
  WG --> HOME["index.md via README pipeline"]
  WG --> PAGES["6 pages in ThreadPoolExecutor"]
  WG --> API["api/ pages per filtered symbol"]
  PAGES --> RAG
  PAGES --> MC["ModelClient (OpenAI / Azure OpenAI)"]
  WG --> MK["mkdocs.yml (fixed nav)"]
Component Path Role
CLI src/readmex/utils/cli.py Mode selection, interactive path prompt, mkdocs serve / deploy helpers
README generator src/readmex/core.py readmex.generate: collect info, summarise files, fill the template
Website generator src/readmex/website_core.py WebsiteGenerator: project analysis, page prompts, API docs, draw.io, mkdocs.yml
Code RAG src/readmex/code_rag.py CodeRAG: Python AST code blocks, embeddings, FAISS search, prompt augmentation
Model client src/readmex/utils/model_client.py Single client for chat and text-to-image, OpenAI or Azure
Config src/readmex/config.py ~/.readmex/config.json + env vars, defaults, ignore and file patterns
Analysers src/readmex/utils/language_analyzer.py, dependency_analyzer.py Language distribution and dependency scanning
Template src/readmex/templates/BLANK_README.md README skeleton with badge and link placeholders

How a request flows

README mode, readmex ./my-project:

  1. Entry. main validates config (prompting for missing keys), resolves the path, and calls readmex(...).generate() (cli.py).
  2. Collect. generate loads config, asks for or defaults the output directory (always a readmex_output/ subfolder), and reads git info, user info, project metadata and the language distribution (core.py, L306-L330).
  3. Summarise files. _get_script_descriptions builds file patterns from the detected primary language’s extensions, falling back to *.py, *.sh and *.ipynb, and adds *.md and *.txt. It sends each whole file to the LLM in a thread pool (10 workers by default) and saves script_descriptions.json (core.py).
  4. Fill gaps. If the user left the fields empty, separate LLM calls produce the project description, the entry file, the key features and the extra info. In debug mode these are placeholders (core.py).
  5. Logo. Unless skipped, generate_logo asks the LLM for a short logo description and then calls the image model (logo_generator.py).
  6. Write. _generate_readme_content substitutes known placeholders in BLANK_README.md, drops lines whose values are empty, and sends the template with all collected context to the LLM in one prompt. It then strips code fences and appends any missing badge reference links (core.py, L1330-L1382).

Website mode, readmex . --website:

  1. generate_website creates website/, runs _analyze_project (file tree, dependencies, AST functions and classes, modules, entry points, git info) and, if RAG is on, extracts code blocks and builds embeddings (website_core.py, L286-L327).
  2. The home page reruns the whole README pipeline in silent mode (L869-L900).
  3. API pages are generated per filtered symbol. The six other pages then run in parallel (L208-L255, L347-L388).
  4. Each page goes through _generate_page_content. It builds the page-specific prompt, augments it with up to eight RAG code blocks when RAG succeeded, and calls the model (L656-L713).
  5. _create_mkdocs_config writes mkdocs.yml with a fixed navigation and the Material theme set to Chinese (L2174-L2200).

Key components

File selection

find_files walks the tree with fnmatch. It combines a built-in ignore list with the project’s .gitignore patterns (file_handler.py). The built-in list includes *.md, every __init__.py down to two levels, and *output*, which matches any path containing “output” (config.py). Because the ignore check runs first, the *.md entry in DOCUMENT_PATTERNS never matches, so existing Markdown docs are not read. Files are sent whole, with no truncation or chunking.

CodeRAG

CodeRAG is used only in website mode and covers only Python. extract_code_blocks runs rglob('*.py') over the project, which ignores .gitignore, and parses each file with ast into function, class, method and import blocks with signature, docstring, complexity and dependencies (code_rag.py). Embeddings come from a local sentence-transformers model by default, Kwaipilot/OASIS-code-embedding-1.5B, or from an OpenAI-compatible embeddings API (L92-L135). Search uses a FAISS inner-product index over normalised vectors, with a keyword fallback. Each page’s query is a fixed Chinese keyword string such as “项目 X 架构 设计 模块结构” for the architecture page. Blocks, relations, embeddings and the index are pickled under website/.rag_cache/. If the cache files exist, they are loaded without any staleness check (L868-L933).

Architecture page

_generate_architecture_page runs two jobs in parallel: page text, and a draw.io XML diagram. The diagram job checks the XML for truncation, retries, and falls back to a default diagram. The diagram is written to architecture_diagram.drawio and embedded as an image reference that the MkDocs drawio plugin renders (website_core.py, L1409-L1515). There are no Mermaid diagrams.

ModelClient

One class wraps both chat and image calls. A base URL containing .openai.azure.com selects AzureOpenAI; anything else uses OpenAI(base_url=...), so any compatible endpoint works. get_answer sends one user message, retries three times with exponential backoff, and uses the constructor’s max_tokens=10000 and temperature=0.7 (model_client.py, L188-L247). The llm_max_tokens and llm_temperature config values exist (config.py), but this client does not use them. All stages share one model.

Extending it

  • Prompts. Each site page has its own _create_<page>_prompt method in website_core.py. Edit them there; they are plain f-strings.
  • Pages. Adding a page means a new prompt method, a generator method, an entry in _generate_pages_in_parallel, and a nav line in _create_mkdocs_config.
  • Models. Set LLM_*, T2I_* and EMBEDDING_* keys in ~/.readmex/config.json or the environment. LOCAL_EMBEDDING=false switches RAG to the embeddings API.
  • Template. The README layout is templates/BLANK_README.md, with {{placeholder}} fields.

Running it

  • Install. pip install readmex. The dependencies include sentence-transformers and faiss-cpu.
  • Configure. On first run, validate_config asks for any missing keys. Keys are read from ~/.readmex/config.json, overridden by environment variables.
  • Run. readmex [path] writes readmex_output/README.md under the chosen output directory. --website writes <project>/website/, and --serve starts mkdocs serve on port 8000-8009. If MkDocs is missing, --serve tries to install it with a hard-coded pip3.9 (cli.py). --debug skips all LLM calls, and --silent skips prompts.

Strengths and caveats

  • Strength: practical README output. For a small project, the template, auto-filled fields, badges and an optional logo give a presentable README in one command.
  • Strength: changelog grounded in git. The changelog prompt is built from git log --stat, commit-type grouping and contributors, not from guesswork.
  • Strength: simple provider story. Any OpenAI-compatible or Azure endpoint, with separate settings for chat, image and embedding models.
  • Caveat: fixed structure, Chinese site. The site always has the same eight sections with Chinese labels, whatever the project is.
  • Caveat: cost and context. README mode makes one LLM call per matched file with the whole file inlined, so large files can overflow context and large repos cost a call each.
  • Caveat: stale RAG. The .rag_cache is never invalidated, so later runs search an old snapshot until you delete the folder. Building it the first time downloads a 1.5B-parameter embedding model by default.
  • Caveat: Python bias. AST extraction, the API pages and RAG all handle only .py files. Other languages get only the per-file summaries.
  • Caveat: rough edges. The CLI reports version 0.1.8 while the package is 0.3.0, logs are a mix of Chinese and English, and configured token limits are ignored.

Sources: code at b83df61, verified Q&A.

How it answers the Open-source DeepWiki questions

Each answer was drafted by a code-reading agent at commit b83df61. Its citations were checked mechanically. Compare with the other open-source deepwiki →

How is a repository ingested and chunked?

answered

The project does not clone repositories — it works exclusively on a local filesystem path provided via CLI arg or interactive input (cli.py:30-34, core.py:267-301). File discovery uses find_files() with glob patterns: SCRIPT_PATTERNS = ['*.py', '*.sh', '*.ipynb'] and DOCUMENT_PATTERNS = ['*.md', '*.txt'] (config.py:235-236). Gitignore patterns are loaded from the project’s .gitignore and merged with a built-in ignore list (DEFAULT_IGNORE_PATTERNS in config.py:211-232) that excludes .git, __pycache__, node_modules, *.md (so markdown files themselves are not re-processed), build/, dist/, virtual envs, IDE dirs, and __init__.py files. A much larger ignore-patterns config lives in config/ignore_patterns.json for the LanguageAnalyzer — hundreds of patterns covering Python, JS, Rust, and more. There is no chunking in the traditional RAG sense: for README generation, each matched file is read in full and its content sent in a prompt to the LLM for a description, using ThreadPoolExecutor (configurable max_workers, default 10) via _get_script_descriptions() (core.py:690-832). For the website path, the CodeRAG.extract_code_blocks() method (code_rag.py:188-217) does AST-level parsing of .py files only — it walks each file’s AST with ast.walk() and extracts FunctionDef, ClassDef, methods, and import statements into individual CodeBlock dataclass instances with type, signature, docstring, complexity score, and dependency lists. No large-repo limits are enforced beyond available memory and thread count.

Editor's note. Correction: README-mode file selection uses the extensions of the primary language detected by LanguageAnalyzer plus *.md/*.txt; SCRIPT_PATTERNS (*.py, *.sh, *.ipynb) is only the fallback. Because *.md is also in the default ignore list, Markdown files are never actually read.

How is retrieval (RAG) implemented?

answered

Retrieval is implemented via the CodeRAG class (code_rag.py:92-948). It has two vectorization paths configured by local_embedding (default true in config.py:197). Local path: loads a SentenceTransformer model, defaulting to Kwaipilot/OASIS-code-embedding-1.5B (code_rag.py:115). Web path: uses the OpenAI embeddings API (text-embedding-3-small) with configurable base_url/api_key (code_rag.py:591-637), sending texts in batches of 100. In either case, embedding text is built from Type, Name, Signature, docstring, code preview (first 500 chars), and dependencies (code_rag.py:526-552). The vectors are stored in a FAISS IndexFlatIP (inner-product similarity) index (code_rag.py:639-649), with L2 normalization applied. semantic_search(query, top_k=5, min_score=0.3) (code_rag.py:651-662) dispatches to either _vector_search() (FAISS or numpy dot-product) or _text_search() (term-matching fallback with weighted keyword scoring). The generate_enhanced_prompt() method (code_rag.py:789-828) wraps retrieved CodeBlock results — type, name, file path, signature, docstring, code snippet (truncated to 300 chars), and dependencies — into a === 相关代码上下文 === section appended to the base page-generation prompt. This is consumed by WebsiteGenerator._generate_page_content() (website_core.py:672-687), which calls code_rag.generate_enhanced_prompt() with a page-specific RAG query (e.g. '项目 {name} 架构 设计 模块结构 类关系 组件' for the architecture page). All code blocks, relations, embeddings, and the FAISS index are cached to pickle and .bin files under .rag_cache/ and loaded from there on subsequent runs (code_rag.py:868-933).

How is the wiki structure (table of contents) determined?

answered

The wiki structure is not LLM-driven — it is hardcoded. WebsiteGenerator._create_mkdocs_config() (website_core.py:2174-2251) produces an mkdocs.yml with a fixed navigation array:

nav:
  - {首页: 'index.md'}
  - {安装: 'installation.md'}
  - {使用说明: 'usage.md'}
  - {API文档: 'api/index.md'}
  - {示例: 'examples.md'}
  - {架构: 'architecture.md'}
  - {贡献指南: 'contributing.md'}
  - {更新日志: 'changelog.md'}

This 8-page structure is always the same regardless of the project. The API docs sub-section expands dynamically because _generate_api_documentation() (website_core.py:347-377) generates one .md file per valuable function/class (filtered by APIDocumentationFilter which skips tests, private symbols, trivial helpers, and low-value files), each saved as api/{module}/{name}.md. The input to the generator is the project_analysis dict built by _analyze_project() (website_core.py:286-327), which contains structure (text tree), functions, classes, modules, dependencies, entry_points, and git_info — all extracted via filesystem traversal and AST parsing. The static nature is deliberate: the progress tracker pre-defines self.stages with those exact 10 stages (website_core.py:38-49) and _generate_pages_in_parallel() (website_core.py:208-255) maps each stage to its generator function explicitly.

How are individual pages generated?

answered

Each page type has a dedicated prompt-builder method (website_core.py:1620-2170): _create_installation_prompt, _create_usage_prompt, _create_examples_prompt, _create_architecture_prompt, _create_contributing_prompt, _create_changelog_prompt. These methods receive the project_analysis dict (file tree, AST-extracted functions/classes, dependency lists, git info, commit history) and format it into a detailed markdown prompt. For the home page, _generate_readme_as_homepage() (website_core.py:869-952) reuses the full README-generation pipeline from readmex.core, creating a temporary readmex instance in silent mode. The architecture page generates a draw.io XML diagram via _generate_drawio_diagram() (website_core.py:1409-1515) with LLM retry logic (3 attempts, validation of XML completeness). The changelog page is especially rich: it runs git log --stat to get commit history with file-change counts, groups commits by type (feat/fix/docs/etc.), detects breaking changes, lists contributors, and feeds all of that into the prompt (website_core.py:2036-2171). Pages are generated in parallel using ThreadPoolExecutor (_generate_pages_in_parallel(), website_core.py:208-255) — six page tasks plus the API doc pages (which themselves use ThreadPoolExecutor for individual API pages). When RAG is enabled, the base prompt is augmented with code_rag.generate_enhanced_prompt() (website_core.py:672-687, code_rag.py:789-828) which appends the top-8 semantically related code blocks as context. Caching: CodeRAG caches extracted code_blocks.pkl, relations.pkl, embeddings.pkl, and faiss_index.bin under .rag_cache/ (code_rag.py:868-933). There is no per-page cache — the RAG cache is all-or-nothing. In debug mode, all LLM calls are skipped and placeholder content is returned instead (website_core.py:715-838). No Mermaid diagrams are generated (the project uses draw.io XML for architecture and MkDocs's drawio plugin for rendering).

How are model providers configured?

answered

The single ModelClient class (utils/model_client.py:10-455) handles both LLM and text-to-image (T2I) providers. It supports Standard OpenAI and Azure OpenAI, detected by checking if the base URL contains .openai.azure.com (model_client.py:60-73). The LLM and T2I providers can differ independently — each has its own config block (get_llm_config() and get_t2i_config() in config.py:168-196). Configuration comes from three sources merged by load_config() (config.py:21-98): (1) ~/.readmex/config.json — a JSON file with keys like LLM_API_KEY, LLM_BASE_URL, LLM_MODEL_NAME (default gpt-3.5-turbo), T2I_MODEL_NAME (default dall-e-3); (2) environment variables (LLM_API_KEY, LLM_BASE_URL, etc.) which override the config file; and (3) interactive prompt from validate_config() if required keys are missing (config.py:100-157). For Azure, the _extract_azure_info() method (model_client.py:75-110) parses the Azure endpoint URL to extract azure_endpoint, deployment_name, and api_version, then creates an AzureOpenAI client. For standard OpenAI, it creates OpenAI(base_url=..., api_key=...) — which means any OpenAI-compatible endpoint works (local models via Ollama/vLLM, Doubao, etc.) by pointing base_url at that server. Per-stage model selection: the same ModelClient instance is used for everything — README generation, page content, API docs, architecture diagrams — so there is no per-stage model routing. However, get_answer() accepts an optional model override parameter (model_client.py:188), and _generate_drawio_diagram() controls quality only through retry logic, not model choice. Embeddings use a separate config (EMBEDDING_API_KEY, EMBEDDING_BASE_URL, EMBEDDING_MODEL_NAME), additionally supporting a LOCAL_EMBEDDING boolean to switch between sentence-transformers and the web API (config.py:190-198).

How is interactive Q&A / chat implemented?

not applicable

This repository does not implement any interactive Q&A, chat, or conversation features. It is a pure batch documentation generator — it reads a project, generates README.md or an MkDocs website, and exits. There is no streaming, no chat loop, no conversation memory, and no MCP exposure. The code_rag.py module's semantic_search() and generate_enhanced_prompt() are used exclusively to augment LLM prompts during page generation, not for end-user querying. The ModelClient.get_answer() method is a simple one-shot LLM call with retry logic and no conversation history. A typical Q&A feature would require a different architecture entirely (interactive loop, session state, web UI), none of which is present.