# How are documents parsed and chunked?

> RAG engines — a good answer covers: Supported formats; OCR and layout parsing; table handling; chunking strategy and sizes.

Canonical page: https://llms-technical-reviews.com/rag/q/ingestion/

## Verdict

[RAGFlow](/p/ragflow/) parses scanned, mixed-format and table-heavy files best. [RAG-Anything](/p/rag-anything/) is the one to use when figures and equations carry the meaning. [Onyx](/p/onyx/) is strongest at pulling content out of workplace apps rather than parsing it.

**Layout-aware parsing.** RAGFlow runs its own ONNX OCR, layout and table-structure models (DeepDoc) and offers twelve chunking templates. The General template cuts at 512 tokens on sentence delimiters. [Kotaemon](/p/kotaemon/) lets each user switch the PDF loader in settings (Adobe, Azure Document Intelligence, Docling, two PaddleOCR modes). It splits text at 1,024 tokens with 256 overlap and keeps tables and figures whole. RAG-Anything parses with MinerU by default. A vision model describes each image, table and equation, and that description becomes a graph entity. LightRAG chunks the plain text.

**Connector-fed apps with plain text extraction.** Onyx has about 50 source connectors. It reads PDFs with pypdfium2 and chunks them into 512-token sentence chunks with no overlap. [AnythingLLM](/p/anything-llm/) splits on characters (1,000 per chunk, 20 overlap). It runs Tesseract OCR only on PDFs that yield no text, and on images, which it OCRs rather than captions. [R2R](/p/r2r/) maps about 35 file types and defaults to 1,024-character chunks with 512 overlap. Only the `recursive` and `character` splitters work. `basic` and `by_title` raise `NotImplementedError`.

**Framework building blocks.** In [LlamaIndex](/p/llama_index/) the default `SentenceSplitter` makes 1,024-token chunks with 200 overlap, and semantic and code splitters are also available. OCR comes only from integrations. [Haystack](/p/haystack/)'s `DocumentSplitter` defaults to 200 words with no overlap, and it ships no OCR engine.

**No classic chunking.** [PageIndex](/p/pageindex/) builds a section tree from pdfium font statistics. Local mode accepts only PDFs and has no OCR. [Quivr](/p/quivr/) cuts 384-token windows with 48 overlap. Its first-party PDF normalizer extracts text one page at a time with no OCR, and it has no Office support.

Pick: RAGFlow for scans, tables and mixed formats.
Pick: RAG-Anything when charts, figures and formulas must be answerable.
Pick: PageIndex for long, well-structured digital PDFs where section boundaries matter more than chunks.

## Per-project answers

### infiniflow/ragflow (answered)

**Formats:** RAGFlow supports PDF, DOCX, TXT, Markdown, HTML, spreadsheets, slides, images, audio, email, and EPUB. Parser type is selected per-document via a `ParserConfig.parser_id` field, with known types including `general`, `naive`, `book`, `paper`, `laws`, `manual`, `presentation`, `qa`, `table`, `picture`, `one`, `audio`, `email`, and `knowledge_graph` (`entity/dataset.go:48-62`).

**PDF handling:** The PDF parser (`deedoc/parser/pdf/parser.go`) is the most sophisticated. It runs a multi-stage pipeline: per-page render → text extraction via both embedded PDF characters and OCR (`parser_ocr.go` uses a DLA detection model + ONNX text recognizer at `recBatchNum=16`), → DLA (Document Layout Analysis) region detection → table extraction (`grid_structure_test.go`) → section assembly. The system can classify text boxes into layout types (title, text, figure, table) and handles de-skewing, rotation candidates, and coordinate-space normalization for all downstream layout algorithms.

**Office documents:** The DOCX parser (`deedoc/parser/office/parser.go`) converts raw blocks to Sections: headings get `layoutType="title"`, tables get `DocTypeKwd="table"` with an HTML row representation, and images get `DocTypeKwd="image"`. The Go port uses `office_oxide` for raw block extraction.

**OCR pipeline:** The OCR pipeline (`parser_ocr.go`) detects text boxes via an ONNX inference model, de-skews each box using `WarpCrop`, generates rotation candidates (0°, ±90°), and runs batched recognition (`recBatchNum=16`). Coordinate scaling divides by zoom factor so downstream layout receives PDF-point coordinates.

**Chunking strategy:** After parsing, the chunker (`ingestion/component/chunker/`) has multiple strategies: token-based (`token.go`), delimiter-based (`delimiter_case_sensitive_test.go`), page-based (`page.go`), section-based (`positions_slice.go`), QA pair (`qa.go`), table-aware (`table.go`, `table_split.go`), and hierarchy-aware (`hierarchy.go`). The `general.go` chunker merges text segments from parsed sections with configurable token caps. Chunk sizes are controlled by parser configuration rather than fixed lengths. The tokenizer component (`tokenizer.go`) computes per-chunk `content_ltks` (tokenized string for BM25) and `content_sm_ltks` (fine-grained variant), plus embedding vectors via the tenant's embedding model.

> **Editor's note.** Correction: the PDF and Office parsers live under `internal/deepdoc/parser/` (not `deedoc/`). `grid_structure_test.go` and `delimiter_case_sensitive_test.go` are tests, not the table or delimiter implementations. The General template chunks at `chunk_token_size: 512` (`internal/ingestion/pipeline/template/ingestion_pipeline_general.json`).

Citations: [internal/entity/dataset.go:44-62](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/entity/dataset.go#L44-L62) · [internal/ingestion/component/tokenizer.go:1-60](https://github.com/infiniflow/ragflow/blob/cc72ecb0af18ade5d58f84d107af7cd59a88c39f/internal/ingestion/component/tokenizer.go#L1-L60)

### Mintplex-Labs/anything-llm (answered)

**Supported formats.** A separate `collector` service (Express.js on port 8888) handles all document ingestion, invoked by the main server via `CollectorApi`. The file-type registry at `collector/utils/constants.js` maps extensions to converter scripts. Text types (`.txt`, `.md`, `.org`, `.adoc`, `.rst`, `.csv`, `.json`) use a simple text extractor. Binary documents: `.pdf` uses PDFLoader with per-page splitting; `.docx`, `.pptx`, `.odt`, `.odp` use office-mime converters (mammoth/docx2js); `.xlsx` is parsed via asXlsx; `.epub` via asEPub; `.mbox` email archives are supported. Audio (`.mp3`, `.wav`, `.m4a`, `.ogg`, `.opus`, `.webm`, `.mp4`) is transcribed to text via whisper (local or cloud). Images (`.png`, `.jpg`, `.webp`) are captioned. YouTube transcripts, website-depth scraping, Confluence/DrupalWiki/Obsidian/Paperless-NGX integrations exist as extension endpoints.

**OCR and layout.** PDF parsing tries text-extraction first via `PDFLoader` (`splitPages: true`). If no text content results, it falls back to `OCRLoader` (Tesseract-based with configurable target languages via `TARGET_OCR_LANG`). Tables are not explicitly detected or extracted — PDF text is concatenated page-wise and no table-structure preservation logic was found. Markdown or HTML tables arrive as raw text.

**Chunking.** The `TextSplitter` class wraps LangChain's `RecursiveCharacterTextSplitter`. Default chunk size is 1,000 characters (configurable via `text_splitter_chunk_size` in SystemSettings, capped at the embedder's `embeddingMaxChunkLength`). Default overlap is 20 characters (configurable via `text_splitter_chunk_overlap`). A `chunkPrefix` can be prepended per embedder requirement (e.g., `search_document: ` for nomic), and a `chunkHeaderMeta` block is prepended with document title, published date, and source URL — wrapped in `<document_metadata>...</document_metadata>` tags so the LLM can cite sources.

> **Editor's note.** Correction: images are not captioned. asImage.js runs Tesseract OCR via OCRLoader.ocrImage and stores the recognised text.

Citations: [collector/utils/constants.js:45-86](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/utils/constants.js#L45-L86) · [collector/processSingleFile/index.js:24-91](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/processSingleFile/index.js#L24-L91) · [collector/processSingleFile/convert/asPDF/index.js:12-84](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/collector/processSingleFile/convert/asPDF/index.js#L12-L84) · [server/utils/TextSplitter/index.js:1-218](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/TextSplitter/index.js#L1-L218) · [server/utils/helpers/index.js:621-630](https://github.com/Mintplex-Labs/anything-llm/blob/297808f9451ceb5795ed208588437f3b526a057a/server/utils/helpers/index.js#L621-L630)

### run-llama/llama_index (answered)

Documents enter through **readers** which implement `BaseReader` or `BasePydanticReader` (`llama_index/core/readers/base.py:19-48`). Core file readers (in `llama-index-integrations/readers/file/`) support PDF, DOCX, EPUB, HTML, Markdown, CSV/Excel, IPYNB, PPTX, RTF, XML, images, and video/audio. Table-specific readers (`PandasCSVReader`, `PandasExcelReader`) parse tabular data. The `UnstructuredReader` wraps the unstructured.io library for OCR and layout parsing but is an integration, not built into core. No built-in OCR or table-extraction logic exists in the core library — those require external integrations.

Chunking is done by **node parsers** (subclasses of `NodeParser` at `llama_index/core/node_parser/interface.py:50-68`). The primary chunkers are:
- **`SentenceSplitter`** (`llama_index/core/node_parser/text/sentence.py:34-100`): default chunk_size=1024 tokens, chunk_overlap=200 tokens, prefers sentence/paragraph boundaries, uses regex `[^,.;。？！]+[,.;。？！]?|[,.;。？！]` as secondary splitter.
- **`TokenTextSplitter`** (`llama_index/core/node_parser/text/token.py:22-84`): splits by token count (using a configurable tokenizer), falls back from separator to `
` to character-level splitting.
- **`SemanticSplitterNodeParser`** (`llama_index/core/node_parser/text/semantic_splitter.py:35-78`): groups sentences by embedding similarity — embeds sentence windows and breaks at a configurable percentile threshold of cosine dissimilarity.
- **`CodeSplitter`** (`llama_index/core/node_parser/text/code.py:19-30`): language-aware AST-based splitting for code files.

Chunking runs in the **ingestion pipeline** (`llama_index/core/ingestion/pipeline.py:72-112`) where `run_transformations()` applies a sequence of `TransformComponent` instances with caching via `IngestionCache`. The pipeline supports parallelism via `ProcessPoolExecutor`.


Citations: [llama-index-core/llama_index/core/readers/base.py:19-48](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/readers/base.py#L19-L48) · [llama-index-core/llama_index/core/node_parser/interface.py:50-68](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/interface.py#L50-L68) · [llama-index-core/llama_index/core/node_parser/text/sentence.py:34-100](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/text/sentence.py#L34-L100) · [llama-index-core/llama_index/core/node_parser/text/token.py:22-84](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/text/token.py#L22-L84) · [llama-index-core/llama_index/core/node_parser/text/semantic_splitter.py:35-78](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/node_parser/text/semantic_splitter.py#L35-L78) · [llama-index-core/llama_index/core/ingestion/pipeline.py:72-112](https://github.com/run-llama/llama_index/blob/81f0e06f6e61a5659bd409cac4a199e28c2fd85e/llama-index-core/llama_index/core/ingestion/pipeline.py#L72-L112)

### The-Vibe-Company/quivr (answered)

The engine does **zero parsing or chunking itself** — it delegates entirely to external plugins. Documents arrive as `Command` structs (content.go:42-80) containing a `Manifest` of typed `Parts` (text, blob, etc.). A **normalizer plugin** (normalization/normalization.go:1-14) runs first for routed binary blobs, converting them into the canonical Manifest/Part structure (the engine ships the NewsML-G2 normalizer for IPTC news formats as an example). The `core.ingest` plugin (plugins/core-ingest/README.md) handles standard text: it only processes **title** and **body** text Parts, cutting body Parts into **windows of at most 384 tokens** (controlled by profile.json), preferring paragraph/line/sentence boundaries, with **48 token overlap**. A title with no body Part is one segment. Refusal limits: 256 KiB text, 2 titles, 64 Parts, or 256 windows (README.md:21-24). The tokenizer (`tokenizer.py`, `tokenizer.json`) is a helper Python process per plugin process, loaded once and kept for the plugin's lifetime (plugins/core-ingest/main.go:53-67). The `IngestionPlugin.SegmentAndEmbed` interface (processing/plugin.go:21) is the single entry point — the plugin returns segments with per-space vectors, and the engine stores the segmentation and embedding artifacts durably. The engine validates the output via the Contract Runner's rules (pluginhttp/ingestion.go:91). No OCR or layout analysis is present in the engine or first-party plugins.


Citations: [internal/content/content.go:42-97](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/content/content.go#L42-L97) · [internal/normalization/normalization.go:1-14](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/normalization/normalization.go#L1-L14) · [plugins/core-ingest/README.md:1-28](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/plugins/core-ingest/README.md#L1-L28) · [plugins/core-ingest/main.go:60-116](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/plugins/core-ingest/main.go#L60-L116) · [internal/processing/plugin.go:1-40](https://github.com/The-Vibe-Company/quivr/blob/63d76a953fb3e42beb6fb771e86eec0f0d0fa009/internal/processing/plugin.go#L1-L40)

### VectifyAI/PageIndex (answered)

**Document format & text extraction.** Local mode only accepts PDFs (`pageindex/local_api.py:119-122`), validated as file path or `BytesIO` (`pageindex/flash/api.py:49-73`). Text is extracted via **PyPDF2** (`local_api.py:209-216`) for the standard pipeline and via **pypdfium2** for the Flash pipeline (`flash/parser_pdfium_parallel.py`), which runs character-level parsing over PDF content streams (`flash/parser_pdfium_charlevel/`). The Flash pipeline uses a **process pool** for parallel page extraction (≥64 pages; `flash/parser_pdfium_parallel.py:47-48`). Cloud mode handles text, scanned, and image-rich documents with managed OCR — the local mode has no OCR (`client.py:849-851`).

**Layout parsing (Flash pipeline).** The Flash pipeline (`flash/main.py:129-303`) performs LLM-free layout analysis: (1) character-level PDF parsing → spans; (2) clustering spans into lines; (3) column detection (`flash/columns/`); (4) clustering lines into blocks with reading order; (5) classification — headers, footers, watermarks, TOC pages, body paragraphs, captions (`flash/classification/`, `flash/labels/`); (6) title detection (`flash/title/`); (7) heading candidate collection and outline assembly (`flash/outline_assembly/`). Table handling is implicit — tables are detected by the `keyword_tables` module (`flash/classification/keyword_tables.py`) and the `has_table_or_prominent` function during outline assembly (`flash/outline_assembly/assembly.py`). No explicit OCR or table extraction runs locally.

**Chunking.** There is **no chunking**. The statement "No Vector DB, No Chunking" is the project's core design (`README.md:16`). Instead, documents are represented as a **hierarchical tree** of sections with page ranges. The standard pipeline (`page_index_classic.py`) groups pages into "groups" of up to 20k tokens with 1-page overlap (`page_list_to_group_text`, line 516-549), but these are processing batches, not retrieval chunks — they're passed to an LLM to extract the TOC structure, not stored as chunks.

**Table of Contents extraction.** Three strategies exist based on TOC presence (`page_index_classic.py:1127-1165`): (1) TOC with page numbers → extract, compute page offset, assign `physical_index`; (2) TOC without page numbers → extract TOC, then LLM-assign page indices; (3) No TOC → LLM generates a tree from page text using `<physical_index_N>` markers. The Flash pipeline does this without LLM by analyzing layout statistics. Embedded PDF bookmarks can optionally supplement the detected structure (`flash/embedded_toc.py`).

> **Editor's note.** Addition: `run_pageindex.py --md_path` builds trees from Markdown headings (`page_index_md.md_to_tree`). Only the client's local mode is PDF-only. In Flash mode the tree comes from pdfium, but the stored page text always comes from PyPDF2 (`local_api._extract_page_texts`).

Citations: [pageindex/local_api.py:119-122](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_api.py#L119-L122) · [pageindex/flash/main.py:129-303](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/flash/main.py#L129-L303) · [pageindex/page_index_classic.py:516-549](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L516-L549) · [pageindex/page_index_classic.py:1127-1165](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1127-L1165)

### onyx-dot-app/onyx (answered)

**Supported formats:** The system accepts PDF, Word (docx), PowerPoint (pptx), Excel (xlsx, xlsm), plain text (txt, md, json, xml, yaml, csv, tsv, sql, conf, log), email (eml/message/rfc822), epub, HTML, and images (png, jpg, jpeg, webp). MIME type classification is in `backend/onyx/file_processing/file_types.py` (`OnyxMimeTypes`, lines 39–76).

**OCR and layout parsing:** Text extraction uses `extract_file_text.py` (line 1+), which delegates PDF to PyMuPDF (via `markitdown`), Office formats to `python-pptx`, `python-docx`, and `openpyxl`, and optionally routes documents through the **Unstructured.io API** (`backend/onyx/file_processing/unstructured.py`, lines 54–70) for richer layout parsing. Images embedded in documents are extracted and optionally summarized by a vision LLM (`image_summarization.py` / `indexing_pipeline.py:process_image_sections`, lines 851–959).

**Chunking strategy:** The `Chunker` class (`backend/onyx/indexing/chunker.py`, lines 124–189) uses the `chonkie` `SentenceChunker` configured by `DOC_EMBEDDING_CONTEXT_SIZE` tokens (default from `shared_configs/`). It splits on sentence boundaries with zero overlap. Each chunk prepends a title prefix and appends a metadata suffix (key-value pairs as natural language). A `DocumentChunker` orchestrates per-document splitting (`backend/onyx/indexing/chunking/document_chunker.py`). The pipeline also generates optional **mini-chunks** (smaller sub-chunks for multipass retrieval) and **large chunks** (combining `LARGE_CHUNK_RATIO` small chunks via `generate_large_chunks`, lines 111–121).

**Table handling:** Tabular content (CSV, xlsx) is processed by a dedicated tabular section chunker and embedded as text. Tables are extracted as text representations rather than preserving native cell structure.

**Contextual RAG enrichment:** During indexing, two optional LLM passes enrich chunks: a document summary (`USE_DOCUMENT_SUMMARY`) and a per-chunk context description (`USE_CHUNK_SUMMARY`). These are generated in parallel (`add_contextual_summaries`, `indexing_pipeline.py:1104–1143`) and stored alongside the chunk embedding.

> **Editor's note.** Correction: PDF text is extracted with pypdfium2 in an isolated process, with pypdf as the fallback, not PyMuPDF via markitdown (backend/onyx/file_processing/extract_file_text.py).

Citations: [backend/onyx/file_processing/file_types.py:39-76](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/file_processing/file_types.py#L39-L76) · [backend/onyx/indexing/chunker.py:124-189](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/indexing/chunker.py#L124-L189) · [backend/onyx/indexing/chunker.py:111-121](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/indexing/chunker.py#L111-L121) · [backend/onyx/indexing/indexing_pipeline.py:851-959](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/indexing/indexing_pipeline.py#L851-L959) · [backend/onyx/indexing/indexing_pipeline.py:1104-1143](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/indexing/indexing_pipeline.py#L1104-L1143) · [backend/onyx/file_processing/unstructured.py:54-70](https://github.com/onyx-dot-app/onyx/blob/a8d6de78eb15aedf9ecee595f4ed999f81c39fda/backend/onyx/file_processing/unstructured.py#L54-L70)

### deepset-ai/haystack (answered)

**Parsing** – Haystack provides a converter component for each supported format. The `PyPDFToDocument` component (`haystack/components/converters/pypdf.py:50-287`) wraps PyPDF2/PyPDF and offers two extraction modes: `PLAIN` mode extracts text as it appears in the PDF stream, while `LAYOUT` mode (experimental, lines 22–23) preserves the rendered layout by considering font height, spacing weight, and vertical whitespace. For images embedded in PDFs, `PDFToImageContent` (`haystack/components/converters/image/pdf_to_image.py:18-80`) renders PDF pages to `ImageContent` objects that can be sent directly to multimodal LLMs — this is the closest Haystack comes to OCR, though no dedicated OCR engine (e.g., Tesseract) is built in. Tables from spreadsheets are handled by `XLSXToDocument` (`haystack/components/converters/xlsx.py:26-59`), which uses pandas/openpyxl to read Excel sheets and outputs each sheet as a Document in CSV or Markdown format. Other converters include `CSVToDocument`, `DOCXToDocument`, `HTMLToDocument`, `MarkdownToDocument`, `PPTXToDocument`, `TextFileToDocument`, and `JSONConverter` — all declared in `haystack/components/converters/__init__.py:5-25`. A `MultiFileConverter` (`haystack/components/converters/multi_file_converter.py:37-58`) dispatches files by MIME type to the appropriate converter via `FileTypeRouter`.

**Chunking** – The `DocumentSplitter` (`haystack/components/preprocessors/document_splitter.py:28-58`) is the primary chunking component. It supports eight `split_by` modes: `word` (default, 200 tokens), `passage` (double newline), `page` (form-feed character), `period`, `line`, `sentence` (NLTK-based), `token` (tiktoken), and `function` (user-supplied callable). Parameters include `split_length`, `split_overlap` (up to split_length-1), and `split_threshold` (minimum units to avoid orphan chunks). When `split_by="token"`, the encoder defaults to `o200k_base` (the tokenizer for current OpenAI models). The `respect_sentence_boundary` option, when combined with `split_by="word"`, uses NLTK to keep whole sentences together. The `HierarchicalDocumentSplitter`, `PythonCodeSplitter`, and `RecursiveSplitter` are also available in the preprocessors directory.


Citations: [haystack/components/converters/__init__.py:5-44](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/converters/__init__.py#L5-L44) · [haystack/components/converters/xlsx.py:26-59](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/converters/xlsx.py#L26-L59) · [haystack/components/converters/image/pdf_to_image.py:18-80](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/converters/image/pdf_to_image.py#L18-L80) · [haystack/components/converters/multi_file_converter.py:37-58](https://github.com/deepset-ai/haystack/blob/e03f7c66fcfa683be6c2b31858f54e60c792b656/haystack/components/converters/multi_file_converter.py#L37-L58)

### Cinnamon/kotaemon (answered)

Documents are parsed by file-type-specific readers (extractors) defined in `KH_DEFAULT_FILE_EXTRACTORS` (`libs/kotaemon/kotaemon/indices/ingests/files.py`, line 48–64). Supported formats: `.pdf`, `.xlsx`, `.docx`, `.pptx`, `.xls`, `.doc`, `.html`, `.mhtml`, `.png`, `.jpeg`, `.jpg`, `.tiff`, `.tif`, `.txt`, `.md`. The PDF reader can be swapped via `reader_mode` to use Adobe PDF Extract API, Azure AI Document Intelligence, or PaddleOCR modes (`libs/ktem/ktem/index/file/pipelines.py`, lines 677–706). In `paddle-struct` mode, the system uses PPStructureV3 for layout analysis, table extraction, and figure detection; `paddle-vl` uses a VLM for vision-language document parsing (`libs/kotaemon/kotaemon/indices/ingests/files.py`, lines 43–45). Table extraction is supported via these specialized readers — tables are stored with metadata type `"table"` and their HTML originals. Thumbnails are extracted per page for PDFs. After loading, the `IndexPipeline.handle_docs()` method splits text documents using the configured `TokenSplitter` (default: chunk_size=1024 tokens, chunk_overlap=256, separator `

`, backup_separators `[
, ., ​]`) (`libs/ktem/ktem/index/file/pipelines.py`, lines 784–789). A `SentenceWindowSplitter` is also available (`libs/kotaemon/kotaemon/indices/splitters/__init__.py`, lines 31–49). Non-text chunks (tables, images, thumbnails) bypass the splitter and are indexed alongside text chunks. Chunks are batched (200 at a time) into the docstore and vectorstore.


Citations: [libs/kotaemon/kotaemon/indices/ingests/files.py:48-64](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/ingests/files.py#L48-L64) · [libs/kotaemon/kotaemon/indices/ingests/files.py:29-45](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/ingests/files.py#L29-L45) · [libs/ktem/ktem/index/file/pipelines.py:354-421](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/ktem/ktem/index/file/pipelines.py#L354-L421) · [libs/ktem/ktem/index/file/pipelines.py:784-791](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/ktem/ktem/index/file/pipelines.py#L784-L791) · [libs/kotaemon/kotaemon/indices/splitters/__init__.py:10-49](https://github.com/Cinnamon/kotaemon/blob/9ad3e4e49aa35b8acddd235918a5d9753c1cfdf9/libs/kotaemon/kotaemon/indices/splitters/__init__.py#L10-L49)

### HKUDS/RAG-Anything (answered)

**Supported formats.** PDF, Office (.doc/.docx/.ppt/.pptx/.xls/.xlsx), images (PNG, JPEG, BMP, TIFF, GIF, WebP), plain text (.txt), Markdown (.md), HTML, and via optional extras: audio (MP3, WAV, FLAC, M4A, OGG, WMA, AAC, OPUS) and video (MP4, MOV, WebM, AVI, MKV, FLV, WMV, M4V). Office files require a separate LibreOffice install for PDF conversion (`parser.py:134-136`).

**Parsers and OCR/layout.** Three backends — MinerU (default), Docling, and PaddleOCR — selected via `RAGAnythingConfig.parser` (`config.py:29-31`). MinerU runs as a subprocess (`parser.py:1357`), extracting structured content including layout blocks (headers, footnotes), tables, equations, and images. It supports `auto`, `txt`, and `ocr` parse methods. PaddleOCR provides better OCR for scanned PDFs via the `[paddleocr]` extra. MinerU's v2 content list (`mineru_content.py:66-104`) is a nested `list[list[dict]]` (pages of blocks); the `convert_mineru_content_list_v2` function flattens it into a uniform block list, optionally stripping page-layout artifacts via `include_layout_blocks`. The Docling parser uses the Docling Python API directly (`parser.py:2224-2229`), avoiding the JSON disk round-trip.

**Table handling.** Tables are preserved as structured blocks with `table_body` and `table_data` fields processed by the `TableModalProcessor` (`modalprocessors.py`), which uses the LLM to generate captions and entity summaries. The utility `format_table_body` (`utils.py:39-63`) renders list-of-list tables as Markdown for prompts/chunks.

**Chunking.** The system delegates chunking to LightRAG's token-size chunker (`chunking_by_token_size`), configured via `chunk_token_size` and `chunk_overlap_token_size` (`raganything.py:92`). Text content is first extracted from parsed blocks, then inserted into LightRAG via `ainsert`, where LightRAG's own pipeline performs the splitting. The `process_document_complete` method in `processor.py:1961-2153` orchestrates parse → text insert → multimodal process in sequence. Context extraction supports page-based and chunk-based modes (`context_window`, `context_mode` config fields) to provide surrounding context to modal processors.


Citations: [raganything/config.py:29-31](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/config.py#L29-L31) · [raganything/parser.py:134-136](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/parser.py#L134-L136) · [raganything/mineru_content.py:66-104](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/mineru_content.py#L66-L104) · [raganything/processor.py:1961-2100](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/processor.py#L1961-L2100) · [raganything/utils.py:39-63](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/utils.py#L39-L63) · [raganything/raganything.py:85-93](https://github.com/HKUDS/RAG-Anything/blob/8664e8b318a3ed651fe62dfcfddfb1f1d633d5be/raganything/raganything.py#L85-L93)

### SciPhi-AI/R2R (answered)

**Supported formats.** R2R supports 40+ document types via a parser registry at `py/core/providers/ingestion/r2r/base.py:40-76`. DEFAULT_PARSERS maps each `DocumentType` enum (`py/shared/abstractions/document.py:18-93`) to a dedicated parser class: BMPParser, CSVParser, DOCParser, DOCXParser, EMLParser, EPUBParser, HTMLParser, JSONParser, MDParser, MSGParser, ORGParser, BasicPDFParser (plus OCRPDFParser, VLMPDFParser, PDFParserUnstructured as extras), PPTParser, PPTXParser, RTFParser, TextParser, TSVParser, XLSParser, XLSXParser, ImageParser (GIF/JPEG/JPG/PNG/HEIC/SVG/TIFF), AudioParser (MP3), P7SParser, RSTParser, PythonParser, JSParser, TSParser, CSSParser.

**OCR and layout parsing.** The `MistralOCRProvider` (`py/core/providers/ocr/mistral.py:13-53`) wraps the Mistral OCR API for document image processing. For PDFs, alternative parsers include OCRPDFParser (local OCR via pdf2image + Mistral) and VLMPDFParser (zero-shot VLM-based). The `UnstructuredIngestionProvider` (`py/core/providers/ingestion/unstructured/base.py`) integrates with Unstructured's API or a local Unstructured service for layout-aware partitioning with options like `pdf_infer_table_structure`, `hi_res_model_name`, and `ocr_languages`. Table handling is further supported through XLSXParserAdvanced and CSVParserAdvanced.

**Chunking strategy.** The `R2RIngestionProvider` (`py/core/providers/ingestion/r2r/base.py:160-215`) builds a text splitter based on `ChunkingStrategy` (RECURSIVE, CHARACTER, BASIC, BY_TITLE). Default is `RecursiveCharacterTextSplitter` with `chunk_size=1024` and `chunk_overlap=512` (`py/core/providers/ingestion/r2r/base.py:31-35`). The recursive splitter (`py/shared/utils/splitter/text.py:1219-1289`) walks through `["

", "
", " ", ""]` separator levels. The character splitter (`py/shared/utils/splitter/text.py:620-646`) uses a configurable separator. For OCR/VLM parsers, an optional `vlm_ocr_one_page_per_chunk` mode creates one `DocumentChunk` per page rather than re-splitting.

> **Editor's note.** Correction: only `recursive` (default, 1,024 characters with 512 overlap) and `character` splitting are implemented in the R2R provider; `basic` and `by_title` raise NotImplementedError.

Citations: [py/core/providers/ingestion/r2r/base.py:31-76](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/providers/ingestion/r2r/base.py#L31-L76) · [py/core/providers/ingestion/r2r/base.py:160-215](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/providers/ingestion/r2r/base.py#L160-L215) · [py/core/providers/ocr/mistral.py:13-53](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/providers/ocr/mistral.py#L13-L53) · [py/core/providers/ingestion/unstructured/base.py:86-111](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/core/providers/ingestion/unstructured/base.py#L86-L111) · [py/shared/utils/splitter/text.py:1219-1289](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/shared/utils/splitter/text.py#L1219-L1289) · [py/shared/abstractions/document.py:18-93](https://github.com/SciPhi-AI/R2R/blob/9c5a94d151f90876bd7eb860f300a8fd662dc481/py/shared/abstractions/document.py#L18-L93)
