# VectifyAI/PageIndex

> Python SDK that builds a page-ranged section tree per PDF and lets an LLM agent read it with tools instead of vector search.

- Category: [RAG engines](https://llms-technical-reviews.com/rag/)
- Repository: https://github.com/VectifyAI/PageIndex (reviewed at commit `6d23caf416858f2ca136840305d1f479a86f6ef7`, 2026-10-01)
- Stars: 38760 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/pageindex/

## Overview

PageIndex is a Python SDK (`pip install pageindex`) for question answering over long PDFs without embeddings. It does not chunk a document into vectors. It builds one table-of-contents tree per document: a node holds a title, a page range (`start_index`/`end_index`), an optional LLM summary and children. At question time an LLM agent gets tools to list documents, read the tree and fetch page text. The agent decides where to look, the way a person uses a book's contents page.

One client class, `PageIndexClient`, runs in two modes. Without an API key everything is local: indexing calls your LLM through LiteLLM, and documents are JSON files on disk. With a `PAGEINDEX_API_KEY`, indexing, OCR and storage happen in the vendor's hosted service, and the SDK becomes a thin client. You can still choose to run the chat agent with your own model. This review covers the local, open code. Cloud features such as folders, OCR and block-level citations are only HTTP calls here.

Choose PageIndex for structured, long PDFs (filings, manuals, papers) where section headings carry meaning and you want page-level provenance. Do not choose it for large corpora of short, flat text, for scanned PDFs in local mode (there is no OCR), or for latency-sensitive search: every question is a multi-turn agent run.

## Architecture

```mermaid
flowchart LR
  U["Your code"] --> C["PageIndexClient"]
  C -->|"no api_key"| L["LocalAPI"]
  C -->|"api_key"| CL["CloudAPI (REST)"]
  L --> F["Flash: layout tree (pdfium)"]
  L --> S["Standard: LLM TOC (PyPDF2)"]
  F --> O["tree_optimize: merge / expand"]
  O --> SUM["Node summaries (LLM)"]
  S --> SUM
  SUM --> ST["DocStore JSON (.pageindex)"]
  C --> CH["Chat agent (openai-agents / anthropic)"]
  CH --> T["agent_tools"]
  T --> ST
  T -->|"cloud"| MB["McpBridge"]
```

| Component | Path | Role |
|---|---|---|
| Client | `pageindex/client.py` | Mode switch, `submit_document`, `chat`, `chat_completions`, tool and agent adapters, citation parsing |
| Local backend | `pageindex/local_api.py` | Validates the PDF, runs Flash or Standard indexing, saves the result |
| Flash indexer | `pageindex/flash/` | Layout-statistics TOC extraction from pdfium character data, with no LLM |
| Standard indexer | `pageindex/page_index_classic.py` | LLM-driven TOC detection, extraction, verification and node splitting |
| Tree optimizer | `pageindex/tree_optimize.py` | Cost-based merge (deterministic) and expand (LLM) passes |
| Store | `pageindex/local_store.py` | `manifest.json` plus per-document `tree.json` and `pages.json`, written atomically |
| Agent tools | `pageindex/agent_tools.py` | `browse_documents`, `get_document`, `get_document_structure`, `get_page_content`, `remove_document`, plus prompts |
| Chat runtime | `pageindex/local_chat.py` | Runs the agent through the OpenAI Agents SDK (via LiteLLM) or Anthropic's tool runner, and handles streaming |
| Cloud bridge | `pageindex/mcp_bridge.py`, `cloud_api.py` | MCP client for the hosted tool server, and the REST client |
| Integrations | `pageindex/integrations/` | Tool adapters for the OpenAI Agents SDK, the Anthropic SDK and the Claude Agent SDK |

## How a request flows

1. `PageIndexClient(...)` without a cloud key loads `config.yaml` through `ConfigLoader` and builds a `LocalAPI` with `storage_path` defaulting to `./.pageindex` ([client.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/client.py#L593-L670)).
2. `submit_document(path)` rejects anything that is not `.pdf`, reads every page's text with PyPDF2, then indexes in `flash` mode (the default) or `standard` mode ([local_api.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_api.py#L100-L170)).
3. `page_index_flash` calls `extract_toc`. It parses spans with pypdfium2 (in parallel), clusters lines into blocks, detects headers, footers, watermarks and TOC pages, picks the title and assembles an outline. Trustworthy PDF bookmarks can frame the result ([main.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/flash/main.py#L124-L200)). If no hierarchy is found, each page becomes a node.
4. The tree then goes through `tree_optimize` (merge, plus an LLM "expand" when the page text is readable) and bottom-up node summaries, capped at 150 words by default ([api.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/flash/api.py#L180-L248)).
5. The store writes `tree.json` (with node text removed) and `pages.json`, and updates the manifest. The returned `doc_id` is `pi-<uuid>`.
6. `client.chat(messages, doc_id=...)` builds an OpenAI Agents SDK `Agent` named `PageIndex`. It gets the managed instructions, a "targeting block" sent as the first user message, and tools limited to the selected documents ([local_chat.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L335-L402), [local_chat.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L637-L660)).
7. The agent calls `get_document_structure`, which is split into parts when the tree is large, and then `get_page_content("5-7")` ([agent_tools.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/agent_tools.py#L1009-L1040)). It answers from what it read. With `citations=True` it adds `<cite doc=… page=…/>` tags, and `_parse_citations` turns them into structured references ([client.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/client.py#L48-L88)).

## Key components

### Flash indexing

Flash is the new default and the main engineering in the repo. More than 70 modules under `pageindex/flash/` re-implement PDF text extraction at the character level (content streams, CMaps, font Unicode repair) and then infer headings from font statistics, column gutters and reading order. The tree comes from layout alone, so `summary=False, optimize=False` produces a structure with no LLM calls. The stored page text comes from PyPDF2, not pdfium, so `_check_page_bounds` rejects a tree that points past the pages PyPDF2 could read.

### Standard indexing

The original pipeline asks the LLM to find a TOC in the first `toc_check_page_num` (20) pages. It then extracts entries with page numbers, maps them to physical pages and checks a sample of titles against the page text with `verify_toc`. If accuracy is 1.0 the tree is accepted. Above 0.6 the wrong entries are fixed with up to 3 retries. Otherwise it falls back from "TOC with page numbers" to "TOC without page numbers" to "no TOC" ([page_index_classic.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1127-L1166)). Nodes larger than 10 pages and 20,000 tokens are split recursively ([page_index_classic.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1200-L1230)).

### Tree optimization

`tree_optimize.py` counts search cost in pages. Routing through a node costs one page, and scanning a collapsed node costs its whole span. `merge` collapses subtrees that cost more to route than to scan, and keeps the removed titles as `key_items`. `expand` asks the LLM to propose subsections for nodes whose span is too large ([tree_optimize.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/tree_optimize.py#L1-L50)).

### Chat lanes

The answer lane runs `openai-agents` with a `LitellmModel`, so any LiteLLM model works. It sets `RunConfig(tracing_disabled=True)` ([local_chat.py](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L483-L490)). `protocol="responses"` and `protocol="chat_completions"` return native OpenAI envelopes. `protocol="messages"` uses Anthropic's own tool runner. Streaming runs an async generator on a background thread and can weave in thinking, tool calls and tool results.

## Extending it

- **Your own agent:** `client.as_openai_tools()`, `as_anthropic_tools()`, `as_claude_mcp()` (an in-process SDK MCP server in local mode) and `agent_instructions()` let you put PageIndex tools into an existing agent.
- **Models:** `index_model` and `chat_model` are separate LiteLLM names (`anthropic/...`, `bedrock/...`, `vertex_ai/...`). The defaults are `gpt-5.6-luna` and `gpt-5.6-sol`.
- **Index tuning:** use `config.yaml` keys, or `run_pageindex.py` flags such as `--mode`, `--optimize`, `--summary-max-words` and `--embedded-toc`. The CLI also builds trees from Markdown headings (`--md_path`, `page_index_md.md_to_tree`), but the client's local mode accepts only PDFs.

## Running it

Install with `pip install pageindex` (Python 3.10 or later). Add `[anthropic]` or `[claude]` for those SDKs. Set the provider key that LiteLLM expects (for example `OPENAI_API_KEY`). No database, vector store or server is needed: the store is a directory of JSON files guarded by a lock. To index a file without writing code, run `python run_pageindex.py --pdf_path doc.pdf`, which writes the tree JSON to `results/`.

## Strengths and caveats

- **Strength: navigable, auditable retrieval.** Every answer can be traced to the tool calls and page ranges the agent read, and citations are page-exact.
- **Strength: zero infrastructure.** The local mode is a library plus a folder, and Flash can build a tree with no LLM cost.
- **Strength: tests.** Over 12,000 lines of tests cover the client, the tools, chat and Flash extraction.
- **Caveat: cost and latency per question.** Each answer is an agent loop of several LLM turns over long page text. There is no cheap first-pass filter.
- **Caveat: the scope of local mode.** It accepts PDFs only, has no OCR (scanned pages give an empty or page-per-node tree) and cites at page level only. Folders, metadata filters and block citations are cloud-only, and the local tools return an error for them.
- **Caveat: no evaluation or tracing hooks.** Agents SDK tracing is turned off. Benchmarks live in separate repositories.
- **Caveat: a rough legacy path.** `page_index_classic.py` still `print()`s debug output during indexing.

*Sources: code at 6d23caf, verified Q&A.*

## How VectifyAI/PageIndex answers the RAG engines questions

### How are documents parsed and chunked? (answered)

**Document format & text extraction.** Local mode only accepts PDFs (`pageindex/local_api.py:119-122`), validated as file path or `BytesIO` (`pageindex/flash/api.py:49-73`). Text is extracted via **PyPDF2** (`local_api.py:209-216`) for the standard pipeline and via **pypdfium2** for the Flash pipeline (`flash/parser_pdfium_parallel.py`), which runs character-level parsing over PDF content streams (`flash/parser_pdfium_charlevel/`). The Flash pipeline uses a **process pool** for parallel page extraction (≥64 pages; `flash/parser_pdfium_parallel.py:47-48`). Cloud mode handles text, scanned, and image-rich documents with managed OCR — the local mode has no OCR (`client.py:849-851`).

**Layout parsing (Flash pipeline).** The Flash pipeline (`flash/main.py:129-303`) performs LLM-free layout analysis: (1) character-level PDF parsing → spans; (2) clustering spans into lines; (3) column detection (`flash/columns/`); (4) clustering lines into blocks with reading order; (5) classification — headers, footers, watermarks, TOC pages, body paragraphs, captions (`flash/classification/`, `flash/labels/`); (6) title detection (`flash/title/`); (7) heading candidate collection and outline assembly (`flash/outline_assembly/`). Table handling is implicit — tables are detected by the `keyword_tables` module (`flash/classification/keyword_tables.py`) and the `has_table_or_prominent` function during outline assembly (`flash/outline_assembly/assembly.py`). No explicit OCR or table extraction runs locally.

**Chunking.** There is **no chunking**. The statement "No Vector DB, No Chunking" is the project's core design (`README.md:16`). Instead, documents are represented as a **hierarchical tree** of sections with page ranges. The standard pipeline (`page_index_classic.py`) groups pages into "groups" of up to 20k tokens with 1-page overlap (`page_list_to_group_text`, line 516-549), but these are processing batches, not retrieval chunks — they're passed to an LLM to extract the TOC structure, not stored as chunks.

**Table of Contents extraction.** Three strategies exist based on TOC presence (`page_index_classic.py:1127-1165`): (1) TOC with page numbers → extract, compute page offset, assign `physical_index`; (2) TOC without page numbers → extract TOC, then LLM-assign page indices; (3) No TOC → LLM generates a tree from page text using `<physical_index_N>` markers. The Flash pipeline does this without LLM by analyzing layout statistics. Embedded PDF bookmarks can optionally supplement the detected structure (`flash/embedded_toc.py`).

> **Editor's note.** Addition: `run_pageindex.py --md_path` builds trees from Markdown headings (`page_index_md.md_to_tree`). Only the client's local mode is PDF-only. In Flash mode the tree comes from pdfium, but the stored page text always comes from PyPDF2 (`local_api._extract_page_texts`).

Citations: [pageindex/local_api.py:119-122](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_api.py#L119-L122) · [pageindex/flash/main.py:129-303](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/flash/main.py#L129-L303) · [pageindex/page_index_classic.py:516-549](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L516-L549) · [pageindex/page_index_classic.py:1127-1165](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1127-L1165)

### How are embeddings and indexes built and stored? (answered)

**No embeddings, no vector stores.** PageIndex does not compute vector embeddings or use vector databases. The index is a **hierarchical tree of sections** with page ranges stored as JSON. This "tree index" replaces what a vector index would do in traditional RAG.

**Tree structure.** Each node has `{title, node_id, start_index, end_index, (summary,) (text,) nodes}` (`local_api.py:396-414`, `types.py`). The tree is built bottom-up: leaves from layout headings, parent nodes spanning the range of their children. Intro nodes cover pages a parent opens with before its first child (`tree_optimize.py`). Node summaries are LLM-generated (the `summary_model`) using a bottom-up scheduler — leaves from their own page text, parents composed from child summaries plus residual pages (`utils.py:1140-1157`). Small nodes (under ~300 tokens) keep raw text as their summary.

**Standard mode indexing.** The classic pipeline (`page_index_classic.py:1200-1283`): (1) detect TOC pages via LLM, (2) extract the TOC structure with page indices, (3) verify correctness by spot-checking titles against page text (`verify_toc`, line 1066-1120), (4) fix incorrect items with retries (`fix_incorrect_toc_with_retries`, line 1044-1060), (5) post-process into a nested tree, (6) recursively split large nodes (>10 pages and >20k tokens; `process_large_node_recursively`, line 1168-1198), (7) add node text and summaries, (8) generate an optional one-sentence doc description (`generate_doc_description`, `utils.py:1183-1198`). Configurable via `config.yaml`: defaults `toc_check_page_num=20`, `max_page_num_each_node=10`, `max_token_num_each_node=20000`.

**Flash mode indexing.** `flash/api.py:180-248`: (1) LLM-free layout-based TOC extraction (`extract_toc`), (2) optional embedded bookmark integration, (3) deterministic `merge` pass (collapses nodes where navigating to children is more expensive than scanning the parent; `tree_optimize.py:18-24`), (4) optional LLM `expand` pass (proposes subsections where a node's span exceeds cost of routing; `tree_optimize.py:12-16`), (5) bottom-up summaries. The entire Flash pipeline is LLM-free when `summary=False, optimize=False`.

**Local storage.** The local store (`local_store.py`) writes JSON files to `~/.pageindex/`: each document as a directory containing `doc.json` (metadata), `tree.json` (the tree structure), `pages.json` (page text). A `manifest.json` indexes all documents. Atomic writes with `_write_json_atomic` use temp-file + `os.replace`. No external database is required.

**Cloud indexing.** Cloud mode delegates indexing to PageIndex's managed pipeline at `api.pageindex.ai` (`cloud_api.py:38-86`). The client uploads the PDF and receives a `doc_id`; the cloud handles parsing, OCR, image understanding, tree construction, and stores the result. Metadata, folders, and block-level citations are cloud-only features (`README.md:199-208`).

> **Editor's note.** Correction: the local store defaults to `./.pageindex` relative to the working directory (`storage_path`, `client.py`), not `~/.pageindex`.

Citations: [pageindex/local_api.py:396-414](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_api.py#L396-L414) · [pageindex/utils.py:1140-1157](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/utils.py#L1140-L1157) · [pageindex/page_index_classic.py:1200-1283](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1200-L1283) · [pageindex/tree_optimize.py:1-54](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/tree_optimize.py#L1-L54)

### How is retrieval performed? (answered)

**There is no traditional retrieval.** PageIndex does not perform dense/sparse/hybrid search, BM25, query rewriting, decomposition, or reranking. Instead, retrieval is an **agentic LLM reasoning process** over the tree index.

**The agent-based retrieval mechanism.** The chat model receives a system prompt that includes the tree structure and instructions for navigating it (`agent_tools.py:1587-1627`). It uses a set of **MCP-style tools** to explore the document: `browse_documents` (list documents), `get_document` (check status/metadata), `get_document_structure` (fetch the hierarchical tree outline), `get_page_content` (retrieve page text by page numbers), and `remove_document` (`agent_tools.py:68-303`). The `get_document_structure` tool returns the tree outline paginated across multiple parts when large. The agent decides which tool to call, what parameters to use, and what pages to read based on the question — effectively performing its own query routing.

**Document targeting.** For scoped retrieval, the system prepends a **targeting block** as the first user message: the document's metadata and a directive to work within it (`local_chat.py:1695-1777`). For folder-scoped searches (cloud only), a `folder_targeting_block` is added. The targeting block replaces what would be query filtering in a vector system.

**Multi-document search.** Multiple `doc_id` values can be passed to `chat()`. The targeting block lists all documents, and the agent can call `get_document_structure` and `get_page_content` on each. The conversation's scope is enforced at the tool layer via `_allowed_ids` (`agent_tools.py:1238-1239`).

**Filters.** Local mode supports `folder_id`/`sort`/`query` on `browse_documents` but returns an error — those are cloud-only (`agent_tools.py:727-749`). Local metadata is stored and returned but not searchable as a filter. The tree structure's `start_index`/`end_index` page ranges serve as an implicit filter: the agent can target specific page ranges via `get_page_content("5-10")`.

**Cloud differences.** With a cloud API key, the client connects to the PageIndex MCP server (`mcp_bridge.py`) which serves the live tool set including folders. The cloud also provides a managed chat endpoint at `/chat/completions/` (`cloud_api.py:272-340`) that selects its own model and can `enable_citations`. In own-model chat, retrieval runs the same agent loop but over cloud-hosted documents, proxied through the same tool interface.


Citations: [pageindex/agent_tools.py:68-303](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/agent_tools.py#L68-L303) · [pageindex/local_chat.py:637-659](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L637-L659) · [pageindex/agent_tools.py:1587-1627](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/agent_tools.py#L1587-L1627)

### How are answers generated and grounded? (answered)

**Prompt assembly.** Generation uses the OpenAI Agents SDK or Anthropic SDK tool runner (`local_chat.py`). The system prompt is assembled from three parts: `CHAT_HEADER` (`"You are PageIndex by Vectify AI…"`), the `_base_instructions` (tool descriptions and discovery workflow), and the client's `instructions` plus any system/developer messages from the conversation history. The `_managed_instructions` function combines these (`local_chat.py:28-33`). For the chat_completions lane, user messages are validated and system rows are extracted into the instructions (`_split_chat_messages`, line 50-85).

**Document grounding.** A **targeting block** is prepended as the first user message — it lists the document(s) with metadata and directs the agent to use the tools (`local_chat.py:645-647`). The agent must use tools (`get_document_structure`, `get_page_content`) to read from actual documents; it cannot answer from general knowledge. The tool descriptions instruct the agent: "Cite only statements supported by tool outputs" and "Never fill a gap from general knowledge" (`agent_tools.py:1636-1644`).

**Citations.** When `citations=True`, a citation prompt (the `cited_answer` MCP server prompt or its local frozen copy) is appended to instructions, teaching the model to emit `<cite doc="..." page="..."/>` tags after each claim (`agent_tools.py:1634-1655`, `client.py:1391-1402`). The client parses these tags into structured citation objects (`client.py:48-88`). Cloud mode supports block-level citations (`block_id`); local mode is page-level only.

**Streaming.** `chat(stream=True)` returns a `ChatStream` — iterating yields text chunks with the agent's process woven in (thinking, tool calls, results) when `show_process=True` (default). The raw event stream via `.events` yields typed dicts: `thinking`, `answer`, `tool_call`, `tool_result` (`local_chat.py:721-773`). A sync pump drives an async generator on a background thread (`_stream_sync`, line 96-160). The managed endpoint's stream uses its own chunk parsing (`_cloud_chunk_events`, line 845-888).

**Protocol lanes.** Three wire protocols: (1) The default **answer lane** — simplified Chat Completions envelope; (2) `protocol="responses"` — native OpenAI Responses API with full transcript (tool calls/results as native items), enabling prompt-cache continuity across turns; (3) `protocol="messages"` — native Anthropic Messages API with the SDK's tool runner, cache-control markings, and cross-turn aggregated usage. All three support multi-turn via tool_use/tool_result round-trips.

**Agentic answering.** The answering agent uses `max_turns` (default 10) to limit the reasoning loop (`local_chat.py:404-408`). It can iteratively explore the document structure, read specific pages, and refine its understanding. The OpenAI Agents SDK `Runner` handles the turn loop; the Anthropic SDK's own `tool_runner` does the same for the Messages lane.


Citations: [pageindex/local_chat.py:28-33](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L28-L33) · [pageindex/agent_tools.py:1634-1655](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/agent_tools.py#L1634-L1655) · [pageindex/local_chat.py:721-773](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L721-L773) · [pageindex/local_chat.py:637-659](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L637-L659)

### How is quality evaluated or observed? (answered)

**No built-in evaluation framework.** PageIndex has no built-in evaluation harness, no metrics (accuracy, precision, recall, F1, faithfulness), no online evaluation, no A/B testing infrastructure, and no tracing/observability hooks for LLM calls. There is no evaluation module, no eval datasets shipped with the repo, and no logging of retrieval or generation quality.

**What exists instead.** The project's evaluation is conducted through a **separate benchmark repository** at [PageIndex-OSS-Benchmark](https://github.com/VectifyAI/PageIndex-OSS-Benchmark) (`README.md:137-149`), which measures the quickstart setup on 62 lookup questions over 34 PDFs from MMLongBench-Doc-V2. Results (accuracy vs. cost per question) are reported externally, not produced by the SDK itself. FinanceBench results (98.7% accuracy) were achieved by the separate [Mafin2.5](https://github.com/VectifyAI/Mafin2.5-FinanceBench) system (`README.md:163-175`).

**TOC verification.** The standard pipeline does have a **TOC quality check** — it verifies a random sample of TOC entries by checking whether each section title actually appears on its assigned page (`page_index_classic.py:1066-1120`). If accuracy is <60%, it falls back to a more expensive extraction method (TOC-with-numbers → TOC-without-numbers → no-TOC). Items with incorrect page indices are fixed with up to 3 retries (`fix_incorrect_toc_with_retries`, line 1044-1060). This is a correctness gate during indexing, not an evaluation metric.

**Tree optimization metrics.** The `tree_optimize.py` module computes worst-case search cost (in pages) before and after merge/expand passes, reported in the `optimize` key of flash results (`flash/api.py:120-125`). This measures the efficiency of the tree structure for agent navigation, not answer quality.

**Observability.** The codebase uses Python's `logging` module throughout but defines no tracing, spans, or observability exports. The `disable_tracing=True` flag is set in the OpenAI Agents SDK `RunConfig` (`local_chat.py:486-489`), explicitly disabling tracing. No OpenTelemetry, LangSmith, or similar integrations exist.

**Cloud-only features.** The managed cloud chat endpoint returns usage statistics (token counts) and citation objects in responses (`cloud_api.py:340`). No evaluation-specific endpoints are exposed.


Citations: [pageindex/page_index_classic.py:1066-1120](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1066-L1120) · [pageindex/page_index_classic.py:1044-1060](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/page_index_classic.py#L1044-L1060) · [pageindex/tree_optimize.py:1-54](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/tree_optimize.py#L1-L54) · [pageindex/local_chat.py:486-489](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_chat.py#L486-L489)

### How is it deployed and operated? (answered)

**Dual-mode architecture: library and cloud service.** PageIndex is both a **Python library** (`pip install -U pageindex`) and a **cloud API service** (`api.pageindex.ai`). The same `PageIndexClient` switches modes based on whether an API key is provided (`client.py:593-602,623-668`).

**Local mode deployment.** In local mode, everything runs on the user's machine. Dependencies: `openai-agents`, `litellm` (for LLM routing), `pypdfium2` (for PDF parsing), `PyPDF2` (fallback extraction), `Pillow` (imaging). No external infrastructure — no vector DB, no database server. Documents are stored on disk via `DocStore` (`local_store.py`). LLM calls route through LiteLLM to OpenAI, Anthropic, or any OpenAI-compatible endpoint. The user provides their own API keys (`OPENAI_API_KEY`, etc.) (`README.md:87-96`).

**Cloud mode deployment.** With a `PAGEINDEX_API_KEY`, documents are uploaded to PageIndex Cloud (`cloud_api.py:38-86`). The cloud handles all parsing, OCR, indexing, and storage. The SDK client talks REST to `https://api.pageindex.ai` for document management and `/chat/completions` for the managed chat endpoint. For own-model chat over cloud documents, the agent runs in-process using the same tools proxied through an MCP bridge to the cloud (`agent_tools.py:1490-1510`, `mcp_bridge.py`).

**Supported models.** The indexing model (for summaries and expand) and chat model (for answering) are independently configurable. Model names follow LiteLLM's convention: bare names → OpenAI, `anthropic/claude-*` → Anthropic, `bedrock/*` → AWS Bedrock, `vertex_ai/*` → GCP Vertex, `openai/*` → explicit OpenAI routing (`client.py:92-98,112-135`). The default indexing model is `gpt-5.6-luna`; default chat model is `gpt-5.6-sol` (from `config.yaml`).

**APIs.** The client provides several protocol options: default answer lane (simplified Chat Completions), `protocol="chat_completions"` (full Chat Completions envelope), `protocol="responses"` (native OpenAI Responses), and `protocol="messages"` (native Anthropic Messages). It also offers an **MCP server** for Claude integration (`mcp_bridge.py`) and plain **Python function tools** for the OpenAI Agents SDK and Claude Agent SDK (`integrations/openai_agents.py`, `integrations/claude_agent_sdk.py`, `integrations/anthropic_sdk.py`).

**No UI.** The open-source SDK has no bundled UI. The PageIndex App is a separate hosted product at `app.pageindex.ai` (`README.md:36`).

**Scaling.** Local scaling is limited to single-machine disk storage. Cloud scaling supports "millions of documents" through the **PageIndex File System** (pages, not this repo's code; `README.md:35`). Multi-tenancy is cloud-only via API keys.

**Packaging.** Published via PyPI as `pageindex`. Requires Python ≥3.10. Optional extras: `pageindex[claude]` for Claude Agent SDK; `pageindex[anthropic]` for Anthropic SDK; `pageindex[openai]` (empty, keeps the install flag valid) (`pyproject.toml:50-54`).

> **Editor's note.** Correction: `mcp_bridge.py` is an MCP client for the hosted PageIndex MCP server, not an MCP server. In local mode `as_claude_mcp()` returns an in-process Claude Agent SDK MCP server, which needs `pageindex[claude]`. The local store defaults to `./.pageindex`.

Citations: [pageindex/client.py:593-668](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/client.py#L593-L668) · [pageindex/local_store.py:1-30](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/local_store.py#L1-L30) · [pageindex/cloud_api.py:38-86](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pageindex/cloud_api.py#L38-L86) · [pyproject.toml:1-67](https://github.com/VectifyAI/PageIndex/blob/6d23caf416858f2ca136840305d1f479a86f6ef7/pyproject.toml#L1-L67)
