LLMs Technical Reviews

How are documents parsed and chunked?

Supported formats; OCR and layout parsing; table handling; chunking strategy and sizes.

Verdict

RAGFlow parses scanned, mixed-format and table-heavy files best. RAG-Anything is the one to use when figures and equations carry the meaning. Onyx is strongest at pulling content out of workplace apps rather than parsing it.

Layout-aware parsing. RAGFlow runs its own ONNX OCR, layout and table-structure models (DeepDoc) and offers twelve chunking templates. The General template cuts at 512 tokens on sentence delimiters. Kotaemon lets each user switch the PDF loader in settings (Adobe, Azure Document Intelligence, Docling, two PaddleOCR modes). It splits text at 1,024 tokens with 256 overlap and keeps tables and figures whole. RAG-Anything parses with MinerU by default. A vision model describes each image, table and equation, and that description becomes a graph entity. LightRAG chunks the plain text.

Connector-fed apps with plain text extraction. Onyx has about 50 source connectors. It reads PDFs with pypdfium2 and chunks them into 512-token sentence chunks with no overlap. AnythingLLM splits on characters (1,000 per chunk, 20 overlap). It runs Tesseract OCR only on PDFs that yield no text, and on images, which it OCRs rather than captions. R2R maps about 35 file types and defaults to 1,024-character chunks with 512 overlap. Only the recursive and character splitters work. basic and by_title raise NotImplementedError.

Framework building blocks. In LlamaIndex the default SentenceSplitter makes 1,024-token chunks with 200 overlap, and semantic and code splitters are also available. OCR comes only from integrations. Haystack’s DocumentSplitter defaults to 200 words with no overlap, and it ships no OCR engine.

No classic chunking. PageIndex builds a section tree from pdfium font statistics. Local mode accepts only PDFs and has no OCR. Quivr cuts 384-token windows with 48 overlap. Its first-party PDF normalizer extracts text one page at a time with no OCR, and it has no Office support.

Pick: RAGFlow for scans, tables and mixed formats. Pick: RAG-Anything when charts, figures and formulas must be answerable. Pick: PageIndex for long, well-structured digital PDFs where section boundaries matter more than chunks.

Per-project answers

infiniflow/ragflow

answered

Formats: RAGFlow supports PDF, DOCX, TXT, Markdown, HTML, spreadsheets, slides, images, audio, email, and EPUB. Parser type is selected per-document via a ParserConfig.parser_id field, with known types including general, naive, book, paper, laws, manual, presentation, qa, table, picture, one, audio, email, and knowledge_graph (entity/dataset.go:48-62).

PDF handling: The PDF parser (deedoc/parser/pdf/parser.go) is the most sophisticated. It runs a multi-stage pipeline: per-page render → text extraction via both embedded PDF characters and OCR (parser_ocr.go uses a DLA detection model + ONNX text recognizer at recBatchNum=16), → DLA (Document Layout Analysis) region detection → table extraction (grid_structure_test.go) → section assembly. The system can classify text boxes into layout types (title, text, figure, table) and handles de-skewing, rotation candidates, and coordinate-space normalization for all downstream layout algorithms.

Office documents: The DOCX parser (deedoc/parser/office/parser.go) converts raw blocks to Sections: headings get layoutType="title", tables get DocTypeKwd="table" with an HTML row representation, and images get DocTypeKwd="image". The Go port uses office_oxide for raw block extraction.

OCR pipeline: The OCR pipeline (parser_ocr.go) detects text boxes via an ONNX inference model, de-skews each box using WarpCrop, generates rotation candidates (0°, ±90°), and runs batched recognition (recBatchNum=16). Coordinate scaling divides by zoom factor so downstream layout receives PDF-point coordinates.

Chunking strategy: After parsing, the chunker (ingestion/component/chunker/) has multiple strategies: token-based (token.go), delimiter-based (delimiter_case_sensitive_test.go), page-based (page.go), section-based (positions_slice.go), QA pair (qa.go), table-aware (table.go, table_split.go), and hierarchy-aware (hierarchy.go). The general.go chunker merges text segments from parsed sections with configurable token caps. Chunk sizes are controlled by parser configuration rather than fixed lengths. The tokenizer component (tokenizer.go) computes per-chunk content_ltks (tokenized string for BM25) and content_sm_ltks (fine-grained variant), plus embedding vectors via the tenant's embedding model.

Editor's note. Correction: the PDF and Office parsers live under internal/deepdoc/parser/ (not deedoc/). grid_structure_test.go and delimiter_case_sensitive_test.go are tests, not the table or delimiter implementations. The General template chunks at chunk_token_size: 512 (internal/ingestion/pipeline/template/ingestion_pipeline_general.json).

Mintplex-Labs/anything-llm

answered

Supported formats. A separate collector service (Express.js on port 8888) handles all document ingestion, invoked by the main server via CollectorApi. The file-type registry at collector/utils/constants.js maps extensions to converter scripts. Text types (.txt, .md, .org, .adoc, .rst, .csv, .json) use a simple text extractor. Binary documents: .pdf uses PDFLoader with per-page splitting; .docx, .pptx, .odt, .odp use office-mime converters (mammoth/docx2js); .xlsx is parsed via asXlsx; .epub via asEPub; .mbox email archives are supported. Audio (.mp3, .wav, .m4a, .ogg, .opus, .webm, .mp4) is transcribed to text via whisper (local or cloud). Images (.png, .jpg, .webp) are captioned. YouTube transcripts, website-depth scraping, Confluence/DrupalWiki/Obsidian/Paperless-NGX integrations exist as extension endpoints.

OCR and layout. PDF parsing tries text-extraction first via PDFLoader (splitPages: true). If no text content results, it falls back to OCRLoader (Tesseract-based with configurable target languages via TARGET_OCR_LANG). Tables are not explicitly detected or extracted — PDF text is concatenated page-wise and no table-structure preservation logic was found. Markdown or HTML tables arrive as raw text.

Chunking. The TextSplitter class wraps LangChain's RecursiveCharacterTextSplitter. Default chunk size is 1,000 characters (configurable via text_splitter_chunk_size in SystemSettings, capped at the embedder's embeddingMaxChunkLength). Default overlap is 20 characters (configurable via text_splitter_chunk_overlap). A chunkPrefix can be prepended per embedder requirement (e.g., search_document: for nomic), and a chunkHeaderMeta block is prepended with document title, published date, and source URL — wrapped in <document_metadata>...</document_metadata> tags so the LLM can cite sources.

Editor's note. Correction: images are not captioned. asImage.js runs Tesseract OCR via OCRLoader.ocrImage and stores the recognised text.

run-llama/llama_index

answered

Documents enter through readers which implement BaseReader or BasePydanticReader (llama_index/core/readers/base.py:19-48). Core file readers (in llama-index-integrations/readers/file/) support PDF, DOCX, EPUB, HTML, Markdown, CSV/Excel, IPYNB, PPTX, RTF, XML, images, and video/audio. Table-specific readers (PandasCSVReader, PandasExcelReader) parse tabular data. The UnstructuredReader wraps the unstructured.io library for OCR and layout parsing but is an integration, not built into core. No built-in OCR or table-extraction logic exists in the core library — those require external integrations.

Chunking is done by node parsers (subclasses of NodeParser at llama_index/core/node_parser/interface.py:50-68). The primary chunkers are:

  • SentenceSplitter (llama_index/core/node_parser/text/sentence.py:34-100): default chunk_size=1024 tokens, chunk_overlap=200 tokens, prefers sentence/paragraph boundaries, uses regex [^,.;。?!]+[,.;。?!]?|[,.;。?!] as secondary splitter.
  • TokenTextSplitter (llama_index/core/node_parser/text/token.py:22-84): splits by token count (using a configurable tokenizer), falls back from separator to to character-level splitting.
  • SemanticSplitterNodeParser (llama_index/core/node_parser/text/semantic_splitter.py:35-78): groups sentences by embedding similarity — embeds sentence windows and breaks at a configurable percentile threshold of cosine dissimilarity.
  • CodeSplitter (llama_index/core/node_parser/text/code.py:19-30): language-aware AST-based splitting for code files.

Chunking runs in the ingestion pipeline (llama_index/core/ingestion/pipeline.py:72-112) where run_transformations() applies a sequence of TransformComponent instances with caching via IngestionCache. The pipeline supports parallelism via ProcessPoolExecutor.

The-Vibe-Company/quivr

answered

The engine does zero parsing or chunking itself — it delegates entirely to external plugins. Documents arrive as Command structs (content.go:42-80) containing a Manifest of typed Parts (text, blob, etc.). A normalizer plugin (normalization/normalization.go:1-14) runs first for routed binary blobs, converting them into the canonical Manifest/Part structure (the engine ships the NewsML-G2 normalizer for IPTC news formats as an example). The core.ingest plugin (plugins/core-ingest/README.md) handles standard text: it only processes title and body text Parts, cutting body Parts into windows of at most 384 tokens (controlled by profile.json), preferring paragraph/line/sentence boundaries, with 48 token overlap. A title with no body Part is one segment. Refusal limits: 256 KiB text, 2 titles, 64 Parts, or 256 windows (README.md:21-24). The tokenizer (tokenizer.py, tokenizer.json) is a helper Python process per plugin process, loaded once and kept for the plugin's lifetime (plugins/core-ingest/main.go:53-67). The IngestionPlugin.SegmentAndEmbed interface (processing/plugin.go:21) is the single entry point — the plugin returns segments with per-space vectors, and the engine stores the segmentation and embedding artifacts durably. The engine validates the output via the Contract Runner's rules (pluginhttp/ingestion.go:91). No OCR or layout analysis is present in the engine or first-party plugins.

VectifyAI/PageIndex

answered

Document format & text extraction. Local mode only accepts PDFs (pageindex/local_api.py:119-122), validated as file path or BytesIO (pageindex/flash/api.py:49-73). Text is extracted via PyPDF2 (local_api.py:209-216) for the standard pipeline and via pypdfium2 for the Flash pipeline (flash/parser_pdfium_parallel.py), which runs character-level parsing over PDF content streams (flash/parser_pdfium_charlevel/). The Flash pipeline uses a process pool for parallel page extraction (≥64 pages; flash/parser_pdfium_parallel.py:47-48). Cloud mode handles text, scanned, and image-rich documents with managed OCR — the local mode has no OCR (client.py:849-851).

Layout parsing (Flash pipeline). The Flash pipeline (flash/main.py:129-303) performs LLM-free layout analysis: (1) character-level PDF parsing → spans; (2) clustering spans into lines; (3) column detection (flash/columns/); (4) clustering lines into blocks with reading order; (5) classification — headers, footers, watermarks, TOC pages, body paragraphs, captions (flash/classification/, flash/labels/); (6) title detection (flash/title/); (7) heading candidate collection and outline assembly (flash/outline_assembly/). Table handling is implicit — tables are detected by the keyword_tables module (flash/classification/keyword_tables.py) and the has_table_or_prominent function during outline assembly (flash/outline_assembly/assembly.py). No explicit OCR or table extraction runs locally.

Chunking. There is no chunking. The statement "No Vector DB, No Chunking" is the project's core design (README.md:16). Instead, documents are represented as a hierarchical tree of sections with page ranges. The standard pipeline (page_index_classic.py) groups pages into "groups" of up to 20k tokens with 1-page overlap (page_list_to_group_text, line 516-549), but these are processing batches, not retrieval chunks — they're passed to an LLM to extract the TOC structure, not stored as chunks.

Table of Contents extraction. Three strategies exist based on TOC presence (page_index_classic.py:1127-1165): (1) TOC with page numbers → extract, compute page offset, assign physical_index; (2) TOC without page numbers → extract TOC, then LLM-assign page indices; (3) No TOC → LLM generates a tree from page text using <physical_index_N> markers. The Flash pipeline does this without LLM by analyzing layout statistics. Embedded PDF bookmarks can optionally supplement the detected structure (flash/embedded_toc.py).

Editor's note. Addition: run_pageindex.py --md_path builds trees from Markdown headings (page_index_md.md_to_tree). Only the client's local mode is PDF-only. In Flash mode the tree comes from pdfium, but the stored page text always comes from PyPDF2 (local_api._extract_page_texts).

onyx-dot-app/onyx

answered

Supported formats: The system accepts PDF, Word (docx), PowerPoint (pptx), Excel (xlsx, xlsm), plain text (txt, md, json, xml, yaml, csv, tsv, sql, conf, log), email (eml/message/rfc822), epub, HTML, and images (png, jpg, jpeg, webp). MIME type classification is in backend/onyx/file_processing/file_types.py (OnyxMimeTypes, lines 39–76).

OCR and layout parsing: Text extraction uses extract_file_text.py (line 1+), which delegates PDF to PyMuPDF (via markitdown), Office formats to python-pptx, python-docx, and openpyxl, and optionally routes documents through the Unstructured.io API (backend/onyx/file_processing/unstructured.py, lines 54–70) for richer layout parsing. Images embedded in documents are extracted and optionally summarized by a vision LLM (image_summarization.py / indexing_pipeline.py:process_image_sections, lines 851–959).

Chunking strategy: The Chunker class (backend/onyx/indexing/chunker.py, lines 124–189) uses the chonkie SentenceChunker configured by DOC_EMBEDDING_CONTEXT_SIZE tokens (default from shared_configs/). It splits on sentence boundaries with zero overlap. Each chunk prepends a title prefix and appends a metadata suffix (key-value pairs as natural language). A DocumentChunker orchestrates per-document splitting (backend/onyx/indexing/chunking/document_chunker.py). The pipeline also generates optional mini-chunks (smaller sub-chunks for multipass retrieval) and large chunks (combining LARGE_CHUNK_RATIO small chunks via generate_large_chunks, lines 111–121).

Table handling: Tabular content (CSV, xlsx) is processed by a dedicated tabular section chunker and embedded as text. Tables are extracted as text representations rather than preserving native cell structure.

Contextual RAG enrichment: During indexing, two optional LLM passes enrich chunks: a document summary (USE_DOCUMENT_SUMMARY) and a per-chunk context description (USE_CHUNK_SUMMARY). These are generated in parallel (add_contextual_summaries, indexing_pipeline.py:1104–1143) and stored alongside the chunk embedding.

Editor's note. Correction: PDF text is extracted with pypdfium2 in an isolated process, with pypdf as the fallback, not PyMuPDF via markitdown (backend/onyx/file_processing/extract_file_text.py).

deepset-ai/haystack

answered

Parsing – Haystack provides a converter component for each supported format. The PyPDFToDocument component (haystack/components/converters/pypdf.py:50-287) wraps PyPDF2/PyPDF and offers two extraction modes: PLAIN mode extracts text as it appears in the PDF stream, while LAYOUT mode (experimental, lines 22–23) preserves the rendered layout by considering font height, spacing weight, and vertical whitespace. For images embedded in PDFs, PDFToImageContent (haystack/components/converters/image/pdf_to_image.py:18-80) renders PDF pages to ImageContent objects that can be sent directly to multimodal LLMs — this is the closest Haystack comes to OCR, though no dedicated OCR engine (e.g., Tesseract) is built in. Tables from spreadsheets are handled by XLSXToDocument (haystack/components/converters/xlsx.py:26-59), which uses pandas/openpyxl to read Excel sheets and outputs each sheet as a Document in CSV or Markdown format. Other converters include CSVToDocument, DOCXToDocument, HTMLToDocument, MarkdownToDocument, PPTXToDocument, TextFileToDocument, and JSONConverter — all declared in haystack/components/converters/__init__.py:5-25. A MultiFileConverter (haystack/components/converters/multi_file_converter.py:37-58) dispatches files by MIME type to the appropriate converter via FileTypeRouter.

Chunking – The DocumentSplitter (haystack/components/preprocessors/document_splitter.py:28-58) is the primary chunking component. It supports eight split_by modes: word (default, 200 tokens), passage (double newline), page (form-feed character), period, line, sentence (NLTK-based), token (tiktoken), and function (user-supplied callable). Parameters include split_length, split_overlap (up to split_length-1), and split_threshold (minimum units to avoid orphan chunks). When split_by="token", the encoder defaults to o200k_base (the tokenizer for current OpenAI models). The respect_sentence_boundary option, when combined with split_by="word", uses NLTK to keep whole sentences together. The HierarchicalDocumentSplitter, PythonCodeSplitter, and RecursiveSplitter are also available in the preprocessors directory.

Cinnamon/kotaemon

answered

Documents are parsed by file-type-specific readers (extractors) defined in KH_DEFAULT_FILE_EXTRACTORS (libs/kotaemon/kotaemon/indices/ingests/files.py, line 48–64). Supported formats: .pdf, .xlsx, .docx, .pptx, .xls, .doc, .html, .mhtml, .png, .jpeg, .jpg, .tiff, .tif, .txt, .md. The PDF reader can be swapped via reader_mode to use Adobe PDF Extract API, Azure AI Document Intelligence, or PaddleOCR modes (libs/ktem/ktem/index/file/pipelines.py, lines 677–706). In paddle-struct mode, the system uses PPStructureV3 for layout analysis, table extraction, and figure detection; paddle-vl uses a VLM for vision-language document parsing (libs/kotaemon/kotaemon/indices/ingests/files.py, lines 43–45). Table extraction is supported via these specialized readers — tables are stored with metadata type "table" and their HTML originals. Thumbnails are extracted per page for PDFs. After loading, the IndexPipeline.handle_docs() method splits text documents using the configured TokenSplitter (default: chunk_size=1024 tokens, chunk_overlap=256, separator `

, backup_separators [ , ., ​]) (libs/ktem/ktem/index/file/pipelines.py, lines 784–789). A SentenceWindowSplitter is also available (libs/kotaemon/kotaemon/indices/splitters/init.py`, lines 31–49). Non-text chunks (tables, images, thumbnails) bypass the splitter and are indexed alongside text chunks. Chunks are batched (200 at a time) into the docstore and vectorstore.

HKUDS/RAG-Anything

answered

Supported formats. PDF, Office (.doc/.docx/.ppt/.pptx/.xls/.xlsx), images (PNG, JPEG, BMP, TIFF, GIF, WebP), plain text (.txt), Markdown (.md), HTML, and via optional extras: audio (MP3, WAV, FLAC, M4A, OGG, WMA, AAC, OPUS) and video (MP4, MOV, WebM, AVI, MKV, FLV, WMV, M4V). Office files require a separate LibreOffice install for PDF conversion (parser.py:134-136).

Parsers and OCR/layout. Three backends — MinerU (default), Docling, and PaddleOCR — selected via RAGAnythingConfig.parser (config.py:29-31). MinerU runs as a subprocess (parser.py:1357), extracting structured content including layout blocks (headers, footnotes), tables, equations, and images. It supports auto, txt, and ocr parse methods. PaddleOCR provides better OCR for scanned PDFs via the [paddleocr] extra. MinerU's v2 content list (mineru_content.py:66-104) is a nested list[list[dict]] (pages of blocks); the convert_mineru_content_list_v2 function flattens it into a uniform block list, optionally stripping page-layout artifacts via include_layout_blocks. The Docling parser uses the Docling Python API directly (parser.py:2224-2229), avoiding the JSON disk round-trip.

Table handling. Tables are preserved as structured blocks with table_body and table_data fields processed by the TableModalProcessor (modalprocessors.py), which uses the LLM to generate captions and entity summaries. The utility format_table_body (utils.py:39-63) renders list-of-list tables as Markdown for prompts/chunks.

Chunking. The system delegates chunking to LightRAG's token-size chunker (chunking_by_token_size), configured via chunk_token_size and chunk_overlap_token_size (raganything.py:92). Text content is first extracted from parsed blocks, then inserted into LightRAG via ainsert, where LightRAG's own pipeline performs the splitting. The process_document_complete method in processor.py:1961-2153 orchestrates parse → text insert → multimodal process in sequence. Context extraction supports page-based and chunk-based modes (context_window, context_mode config fields) to provide surrounding context to modal processors.

SciPhi-AI/R2R

answered

Supported formats. R2R supports 40+ document types via a parser registry at py/core/providers/ingestion/r2r/base.py:40-76. DEFAULT_PARSERS maps each DocumentType enum (py/shared/abstractions/document.py:18-93) to a dedicated parser class: BMPParser, CSVParser, DOCParser, DOCXParser, EMLParser, EPUBParser, HTMLParser, JSONParser, MDParser, MSGParser, ORGParser, BasicPDFParser (plus OCRPDFParser, VLMPDFParser, PDFParserUnstructured as extras), PPTParser, PPTXParser, RTFParser, TextParser, TSVParser, XLSParser, XLSXParser, ImageParser (GIF/JPEG/JPG/PNG/HEIC/SVG/TIFF), AudioParser (MP3), P7SParser, RSTParser, PythonParser, JSParser, TSParser, CSSParser.

OCR and layout parsing. The MistralOCRProvider (py/core/providers/ocr/mistral.py:13-53) wraps the Mistral OCR API for document image processing. For PDFs, alternative parsers include OCRPDFParser (local OCR via pdf2image + Mistral) and VLMPDFParser (zero-shot VLM-based). The UnstructuredIngestionProvider (py/core/providers/ingestion/unstructured/base.py) integrates with Unstructured's API or a local Unstructured service for layout-aware partitioning with options like pdf_infer_table_structure, hi_res_model_name, and ocr_languages. Table handling is further supported through XLSXParserAdvanced and CSVParserAdvanced.

Chunking strategy. The R2RIngestionProvider (py/core/providers/ingestion/r2r/base.py:160-215) builds a text splitter based on ChunkingStrategy (RECURSIVE, CHARACTER, BASIC, BY_TITLE). Default is RecursiveCharacterTextSplitter with chunk_size=1024 and chunk_overlap=512 (py/core/providers/ingestion/r2r/base.py:31-35). The recursive splitter (py/shared/utils/splitter/text.py:1219-1289) walks through `["

", " ", " ", ""] separator levels. The character splitter (py/shared/utils/splitter/text.py:620-646) uses a configurable separator. For OCR/VLM parsers, an optional vlm_ocr_one_page_per_chunkmode creates oneDocumentChunk` per page rather than re-splitting.

Editor's note. Correction: only recursive (default, 1,024 characters with 512 overlap) and character splitting are implemented in the R2R provider; basic and by_title raise NotImplementedError.

How are embeddings and indexes built and stored? →