HKUDS/RAG-Anything
Multimodal document RAG library on LightRAG: parses PDFs, Office, images, audio and video, and adds VLM-captioned entities to its graph.
Overview
RAG-Anything, from the HKU Data Intelligence Lab (HKUDS), is a Python library for RAG over documents that are not only text: PDFs with figures, tables and formulas, Office files, images, and with optional extras audio and video. It is a layer on top of LightRAG, the same lab’s graph-based RAG engine. RAG-Anything handles parsing and multimodal understanding. LightRAG does the chunking, entity and relation extraction, storage and query modes.
The idea is simple. A parser (MinerU by default, or Docling, PaddleOCR or your own) turns a file into a flat “content list” of typed blocks. Text blocks are joined and inserted into LightRAG as usual. Each image, table, equation, audio clip or video becomes a “modal entity”: a vision model or LLM writes a description and an entity summary, which are stored as a chunk, a graph node and vectors. LightRAG’s entity extraction then runs on that description, and every extracted entity gets a belongs_to edge to the modal entity. So a question about a number in a chart can reach the chart through the graph.
It is a library with no server, UI or multi-tenant layer. You give it an llm_model_func, an optional vision_model_func and an embedding_func, and call process_document_complete and aquery. The pinned requirements.txt holds LightRAG below 1.5 and says why: RAG-Anything has been merged into LightRAG 1.5+, and this line’s multimodal extraction is not compatible with it. Treat this repository as the standalone line for the 1.4.x LightRAG API.
Architecture
flowchart TD
F["File (PDF, Office, image, audio, video)"] --> P["Parser: MinerU / Docling / PaddleOCR / custom"]
P --> CL["Content list (typed blocks + page_idx)"]
CL --> SEP["separate_content_with_page_map"]
SEP --> TXT["Text -> LightRAG.ainsert"]
SEP --> MM["Multimodal items"]
MM --> MP["Modal processors (image, table, equation, audio, video)"]
MP --> LLMV["LLM / VLM descriptions + entity_info"]
LLMV --> KG["LightRAG chunks, entities, belongs_to edges"]
TXT --> KG
Q["aquery(mode='mix')"] --> KG
Q --> VLM["VLM-enhanced answer with retrieved images"]
| Component | Path | Role |
|---|---|---|
| Facade | raganything/raganything.py |
RAGAnything dataclass: wraps or builds a LightRAG, creates modal processors and caches |
| Config | raganything/config.py |
RAGAnythingConfig, every field overridable by environment variable |
| Parsers | raganything/parser.py |
MineruParser, DoclingParser, PaddleOCRParser, register_parser registry, LibreOffice conversion |
| Processing | raganything/processor.py |
parse_document, process_document_complete, the 7-stage multimodal pipeline, page provenance |
| Modal processors | raganything/modalprocessors*.py |
Image, table, equation, generic, audio (faster-whisper) and video (scene detection + keyframes) |
| Query | raganything/query.py |
aquery, aquery_with_multimodal, aquery_vlm_enhanced |
| Batch | raganything/batch.py, batch_parser.py |
Folder processing with a concurrency limit |
| Callbacks | raganything/callbacks.py |
Lifecycle hooks and a MetricsCallback |
How a request flows
Take await rag.process_document_complete("report.pdf"), then await rag.aquery("What drove Q3 margin?"):
- Initialise.
_initialize_lightragchecks that the parser is installed, buildsLightRAG(**params)with yourlightrag_kwargsmerged over the defaults, and opens two extra KV namespaces,parse_cacheandmultimodal_status(raganything.py, L420-L465). Then_initialize_processorsregisters one processor per enabled modality, plus a generic fallback (raganything.py). - Parse.
parse_documentchecks the parse cache (keyed on file content and parser options) and otherwise calls the selected parser (processor.py). MinerU runs as amineru -p <file> -o <dir> -m <method>subprocess (parser.py). Docling runs in-process. - Split.
separate_content_with_page_mapreturns the joined text, the list of multimodal items, and character-to-page intervals (utils.py). - Insert text.
insert_text_contentcallslightrag.ainsert, so LightRAG’s token chunker and entity extraction handle the text (utils.py). Afterwards_annotate_text_chunk_pageswritespage_idxranges onto the resulting chunks (processor.py). - Describe multimodal items.
_process_multimodal_content_batch_type_awareruns each item through its processor under a semaphore sized bymax_parallel_insert. The image processor base64-encodes the file, adds surrounding text as context, and asks the vision model for a JSON description andentity_info(processor.py). - Write to the graph. The descriptions are formatted with per-type chunk templates, stored as LightRAG chunks, and inserted as modal entities. LightRAG’s batch extraction then runs over them,
belongs_torelations are added from each extracted entity to its modal entity with weight 10, and LightRAG’s merge step writes the graph (processor.py). - Query.
aquerydefaults tomode="mix"and passes the call tolightrag.aquerywith aQueryParam(query.py). If a vision model is configured, it goes toaquery_vlm_enhancedinstead. That path asks LightRAG for the prompt only, finds image paths in the retrieved context, checks them against allowed directories, inlines them as base64, and sends one multimodal message to the vision model (query.py).
Key components
Parsers
get_parser returns one of the three built-ins or a class registered with register_parser. Built-in names cannot be overridden (parser.py). MinerU is the default and the heaviest: it is a separate CLI with its own models. Office files are converted to PDF with LibreOffice first. Docling uses the Python DocumentConverter directly and also reads HTML. PaddleOCR is the option for scanned PDFs. MinerU failures are mapped to readable hints by _diagnose_mineru_failure.
Modal entities
A modal entity is written to four LightRAG stores: the description chunk to text_chunks and chunks_vdb, a graph node with the entity type and summary, and an entity vector. Then LightRAG extraction runs on the chunk. When you call a processor directly, _create_entity_and_chunk does this for one item (modalprocessors.py). The document pipeline does the same work in batch, in stages 2 to 6 of _process_multimodal_content_batch_type_aware. This is why multimodal content is searchable in every LightRAG mode. Only the generated description is embedded, never the pixels, so retrieval quality depends on the captioning model.
Context-aware captioning
ContextExtractor gives each processor the text around an item, either by page or by neighbouring chunks, cut to a token limit. Captions such as “Figure 3 shows the revenue split above” then refer to the right section. Window size and mode are config fields.
Audio and video
The audio processor transcribes with faster-whisper and groups segments into token-sized sections. The video processor finds scene boundaries with SceneDetect, describes keyframes with the vision model, transcribes the audio track, and merges both by timestamp. Each section becomes its own chunk under the media entity. Both need optional extras. Without them, these items fall back to the generic processor with a warning.
Idempotence
Documents get content-based ids. doc_status and the multimodal_status cache record text and multimodal completion separately, so a re-run skips finished work. force_multimodal_reprocess exists for the case where you change graph or vector backends and need the modal entities written again.
Extending it
- Your own parser. Subclass
Parser, implementparse_documentandcheck_installation, callregister_parser("marker", MarkerParser), and setparser="marker". - Pre-parsed content.
insert_content_listtakes a content list you built yourself (text, image, table and equation dicts withpage_idx) and skips parsing. - Storage and retrieval knobs. Anything LightRAG accepts goes through
lightrag_kwargs: storage backends (Postgres, Neo4j, Milvus and others),top_k, chunk sizes,rerank_model_func. Or pass a readyLightRAGinstance. - Prompts.
prompt.pyandprompts_zh.pyhold the analysis and chunk templates.prompt_manager.pyswitches between them. - Observability. Subclass
ProcessingCallbackfor parse, insert, multimodal and query events.MetricsCallbackcollects counts and timings.
Running it
- Install.
pip install raganything, with extras such as[image],[text],[paddleocr],[audio],[video]or[all]. MinerU comes in asmineru[core]and downloads its models on first use. Office files need a system LibreOffice. - Models. Any async functions with LightRAG’s signatures. The examples cover OpenAI-compatible APIs, Ollama, LM Studio, vLLM and MiniMax.
- Storage. By default LightRAG writes JSON and file-based stores under
working_dir(./rag_storage). For anything shared, configure real backends throughlightrag_kwargs. - Scale.
max_concurrent_filesfor folders and LightRAG’s async limits are the only levers. There is no queue, worker pool or API server, so wrap it yourself (or use LightRAG’s server) for multi-user use.
Strengths and caveats
- Strength: multimodal content in the graph. Figures, tables and equations become entities with edges to the concepts inside them. They are not just captions in a separate index.
- Strength: pluggable parsing. Three serious parsers plus a registry, a parse cache, and page provenance on text chunks.
- Strength: small and readable. About 16k lines of Python with clear stages. The heavy lifting stays in LightRAG.
- Caveat: indexing is expensive. Each multimodal item costs a VLM or LLM call for the description and then LightRAG extraction calls on top. Large, figure-heavy PDFs mean many model calls.
- Caveat: tied to LightRAG < 1.5. Upstream LightRAG now contains this functionality, and this package pins the older API. New projects should compare both before choosing.
- Caveat: library only. No auth, tenancy or service layer. One
working_diris one knowledge base. - Caveat: no evaluation or citation rendering. File paths travel with chunks, but answers are whatever LightRAG or the VLM returns. Quality checks are left to you.
Sources: code at 8664e8b, verified Q&A.
How it answers the RAG engines questions
Each answer was drafted by a code-reading agent at commit 8664e8b. Its citations were checked mechanically. Compare with the other rag engines →
How are documents parsed and chunked?
answeredSupported formats. PDF, Office (.doc/.docx/.ppt/.pptx/.xls/.xlsx), images (PNG, JPEG, BMP, TIFF, GIF, WebP), plain text (.txt), Markdown (.md), HTML, and via optional extras: audio (MP3, WAV, FLAC, M4A, OGG, WMA, AAC, OPUS) and video (MP4, MOV, WebM, AVI, MKV, FLV, WMV, M4V). Office files require a separate LibreOffice install for PDF conversion (parser.py:134-136).
Parsers and OCR/layout. Three backends — MinerU (default), Docling, and PaddleOCR — selected via RAGAnythingConfig.parser (config.py:29-31). MinerU runs as a subprocess (parser.py:1357), extracting structured content including layout blocks (headers, footnotes), tables, equations, and images. It supports auto, txt, and ocr parse methods. PaddleOCR provides better OCR for scanned PDFs via the [paddleocr] extra. MinerU's v2 content list (mineru_content.py:66-104) is a nested list[list[dict]] (pages of blocks); the convert_mineru_content_list_v2 function flattens it into a uniform block list, optionally stripping page-layout artifacts via include_layout_blocks. The Docling parser uses the Docling Python API directly (parser.py:2224-2229), avoiding the JSON disk round-trip.
Table handling. Tables are preserved as structured blocks with table_body and table_data fields processed by the TableModalProcessor (modalprocessors.py), which uses the LLM to generate captions and entity summaries. The utility format_table_body (utils.py:39-63) renders list-of-list tables as Markdown for prompts/chunks.
Chunking. The system delegates chunking to LightRAG's token-size chunker (chunking_by_token_size), configured via chunk_token_size and chunk_overlap_token_size (raganything.py:92). Text content is first extracted from parsed blocks, then inserted into LightRAG via ainsert, where LightRAG's own pipeline performs the splitting. The process_document_complete method in processor.py:1961-2153 orchestrates parse → text insert → multimodal process in sequence. Context extraction supports page-based and chunk-based modes (context_window, context_mode config fields) to provide surrounding context to modal processors.
How are embeddings and indexes built and stored?
answeredRAG-Anything does not implement its own indexing — it delegates entirely to LightRAG, which is a dependency (lightrag-hku<1.5 in requirements.txt:6). The RAGAnything class wraps a LightRAG instance (raganything.py:69) and passes through lightrag_kwargs for storage backend selection.
Embedding models. The user provides an embedding_func callable (e.g. openai_embed from lightrag.llm.openai) when constructing RAGAnything. This is passed directly to LightRAG(...) (raganything.py:425). No embedding model is bundled; any embedding function compatible with LightRAG's EmbeddingFunc signature works.
Vector stores and graph storage. LightRAG supports pluggable backends for kv_storage, vector_storage, graph_storage, and doc_status_storage, all exposed through lightrag_kwargs (raganything.py:86-96). These default to LightRAG's built-in options (typically file-based JSON/JSONL for development, with production backends like MongoDB, Neo4j, or Milvus available by configuring lightrag_kwargs). The RAG-Anything code additionally creates two KV namespaces: parse_cache for caching parsed content lists and multimodal_status for multimodal completion tracking (raganything.py:447-465).
Hybrid / keyword (BM25) indexes. This is handled entirely by LightRAG's query modes. RAG-Anything surfaces them as the mode parameter (local, global, hybrid, naive, mix, bypass) in aquery() (query.py:129-131). LightRAG internally maintains BM25 indexes alongside vector and graph indexes for hybrid search.
Metadata. Each chunk stores page_idx and page_idx_end via the page-provenance system in processor.py:251-385. File-level metadata (path, content hash, parser config) is tracked in doc_status records (processor.py:128-159). The use_full_path config option (config.py:126-128) controls whether file references use the absolute path or just the basename.
How is retrieval performed?
answeredDense / sparse / hybrid search. All retrieval is performed by LightRAG through its aquery() method. The user selects the mode: local (graph-traversal only), global (entity-level graph search), hybrid (vector + graph), naive (vector-only chunk retrieval), mix (combines local and global graph results), or bypass (response generation without retrieval). These are passed as QueryParam(mode=...) (query.py:188-197). RAG-Anything's aquery() and aquery_vlm_enhanced() methods are thin wrappers around LightRAG's aquery().
Reranking. LightRAG exposes rerank_model_func in its configuration, which can be passed through lightrag_kwargs (raganything.py:95). RAG-Anything does not add its own reranking layer.
Query rewriting or decomposition. RAG-Anything does not implement query rewriting or multi-step decomposition. The query is sent as-is to LightRAG, which performs the configured retrieval mode.
Filters. No explicit filter mechanism exists in RAG-Anything for filtering by metadata at retrieval time. LightRAG's own graph traversal effectively acts as a structural filter (local/global modes constrain the search space). For multimodal queries, aquery_with_multimodal() (query.py:221-382) pre-processes multimodal content (images, tables, equations) using LLM/VLM captioning to produce an enhanced query string before calling aquery(). For VLM-enhanced queries (aquery_vlm_enhanced, query.py:384-463), the system retrieves context from LightRAG, scans for image paths, validates them against safe directories, encodes valid images to base64, and builds a multimodal message payload for the vision model.
Source attribution. The QueryParam object supports passing file_paths for citation; LightRAG tracks which chunks belong to which document via the doc_status system, and chunks maintain their full_doc_id reference (processor.py:303).
hybrid combines the local (entity) and global (relation) graph retrievals, and mix, the default in RAG-Anything's aquery, combines the graph retrieval with vector chunk retrieval.How are answers generated and grounded?
answeredPrompt assembly. Generation is delegated to LightRAG's aquery() method, which takes a system_prompt parameter (query.py:129-131). RAG-Anything passes through any system_prompt the user provides. For VLM-enhanced queries, a hardcoded system prompt is used: "You are a helpful assistant that can analyze both text and image content to provide comprehensive answers." (query.py:817), optionally extended by user-provided prompts. For multimodal aquery_with_multimodal(), the system uses extensive prompt templates defined in prompt.py — separate prompts for IMAGE_ANALYSIS_SYSTEM, TABLE_ANALYSIS_SYSTEM, EQUATION_ANALYSIS_SYSTEM, and GENERIC_ANALYSIS_SYSTEM — to generate structured JSON descriptions (with detailed_description, entity_info fields) that are then prepended to the user query.
Citations / source attribution. The processed chunks and entities are stored by LightRAG with provenance: each text chunk records page_idx and page_idx_end via the page-provenance system (processor.py:251-385). Multimodal chunks are associated with a doc_id and file path. The QueryParam supports document-level source tracking, but RAG-Anything does not implement explicit citation rendering in generated answers.
Streaming. The aquery() method accepts stream in **kwargs (query.py:284), which is passed through to LightRAG's QueryParam. However, RAG-Anything's own callback/response handling does not explicitly support streaming responses — the flag is mainly used to disable caching (use_cache = not kwargs.get("stream", False)).
Agentic or multi-step answering. RAG-Anything does not implement agentic reasoning loops or multi-step answering. It is a single-turn retrieval-and-generate pipeline. The multimodal query pipeline (aquery_with_multimodal) does have a two-step flow (pre-process multimodal content → enhanced query → retrieve → generate) but this is not adaptive or iterative.
How is quality evaluated or observed?
answeredBuilt-in evals. RAG-Anything has no built-in evaluation framework (no RAGAS, no faithfulness/relevance metrics, no answer-grounding checks). The only quantitative measurement is the MetricsCallback (callbacks.py:183-287), which aggregates processing pipeline statistics: documents processed/failed, content blocks parsed, parse/insert/multimodal processing times, and query counts/times. It exposes a summary() method returning a human-readable report and a reset() method. This is strictly operational telemetry, not quality evaluation.
Tracing / observability hooks. The ProcessingCallback base class (callbacks.py:61-180) defines lifecycle hooks dispatched by CallbackManager: on_parse_start/complete/error, on_text_insert_start/complete, on_multimodal_start/item_complete/complete, on_query_start/complete/error, on_document_complete/error, and on_batch_start/complete. The ProcessingEvent dataclass captures timestamps, doc_ids, file paths, stages, and error details. Users can subclass ProcessingCallback to integrate with any observability system (e.g., a custom class could send events to OpenTelemetry, Datadog, or a log aggregator). The CallbackManager also supports optional internal event logging (enable_event_log(true)) and thread-safe callback registration (callbacks.py:298-346).
LLM response caching. Both text and multimodal queries can be cached via LightRAG's llm_response_cache (query.py:297-380), which avoids re-generating answers for identical inputs. Cache keys incorporate content fingerprints (SHA-256 of file content, MD5 of query parameters) to detect changes (query.py:59-126).
There are no eval benchmarks, no retrieval-quality metrics (hit rate, MRR, precision), and no generation quality metrics (faithfulness, helpfulness) in the codebase.
How is it deployed and operated?
answeredLibrary, not a service. RAG-Anything is a Python library (PyPI package raganything) with no built-in server, REST API, or web UI. It does not include FastAPI endpoints, a Gradio interface, or a CLI server. It is designed to be imported and used programmatically within a Python application.
Infrastructure and scaling. Required infrastructure includes: a document parser (MinerU, installed as a system-level CLI tool; Docling as Python API; or PaddleOCR optionally), plus an LLM and embedding model accessible via user-provided callables (any OpenAI-compatible API, Ollama local, or custom functions as shown in examples/). The RAGAnythingConfig exposes only max_concurrent_files (config.py:68-71) for limited parallelism. There is no built-in horizontal scaling, sharding, or multi-tenant isolation — those would need to be handled by the LightRAG storage backends (e.g., external MongoDB/Neo4j).
Multi-tenancy. Not supported natively. Each RAGAnything instance owns one working_dir (config.py:18-19), which maps to one LightRAG workspace. Separate workspaces would require separate instances or manually swapped working_dir paths.
UI and API. None in the repository. The examples/ directory shows usage patterns: raganything_example.py demonstrates end-to-end processing and queries, ollama_integration_example.py shows local deployment with Ollama, and lmstudio_integration_example.py / vllm_integration_example.py show alternative LLM backends. These are scripts, not a web service.
Operating the parser. MinerU requires system-level installation (mineru[core] PyPI package, which installs the mineru CLI) and may need model data downloaded on first run. Office file support requires LibreOffice installed on the host (parser.py:296-460). These are significant operational dependencies for a production deployment. The MineruExecutionError class (parser.py:111-123) includes known-failure diagnosis with actionable remediation hints for common issues like version mismatches.
Environment configuration. All parser and processing options are configurable via environment variables or the RAGAnythingConfig dataclass (config.py:13-171), including WORKING_DIR, PARSE_METHOD, PARSER, ENABLE_IMAGE_PROCESSING, MAX_CONCURRENT_FILES, and more.