LLMs Technical Reviews
Home / RAG engines / kotaemon

Cinnamon/kotaemon

Gradio app for chatting with documents: keyword plus vector retrieval, rerankers, highlighted citations, GraphRAG/LightRAG indexes and ReAct agents.

GitHub ↗★ 26kPythonApache-2.0commit 9ad3e4e · 2026-05-30homepage ↗

Overview

Kotaemon, from Cinnamon, is a self-hosted Gradio web app for asking questions about your own documents. You upload files into “collections”, choose which files to search, and chat. Answers stream in with highlighted evidence in a built-in PDF viewer, an optional mind map, and relevance scores per retrieved chunk. It is aimed at two groups: end users who want a local “chat with PDFs” tool with multi-user login, and developers who want a hackable RAG reference app.

The repository has two Python packages. libs/kotaemon is the component library: loaders, splitters, embeddings, LLM wrappers, vector and document stores, rerankers, QA pipelines and agents. They are built on Cinnamon’s theflow Node/Param composition model, and many wrap LangChain or LlamaIndex classes. libs/ktem is the application: Gradio pages, SQLModel tables in SQLite for users, sources and settings, index management, reasoning pipelines, an MCP server registry, and a pluggy extension protocol. Everything is wired in flowsettings.py, a Python settings module of class paths and dicts.

There is no REST API. The product is the UI, and every pipeline runs inside the Gradio process.

Architecture

flowchart LR
  U["Browser"] --> G["Gradio app (ktem.main.App)"]
  G --> CHAT["ChatPage.chat_fn"]
  CHAT --> R["Reasoning: FullQA / Decompose / ReAct / ReWOO"]
  R --> RET["DocumentRetrievalPipeline per index"]
  RET --> VR["VectorRetrieval (vector + full-text)"]
  VR --> VS["Vector store: Chroma (default)"]
  VR --> DS["Doc store: LanceDB FTS (default)"]
  VR --> RR["Rerankers: Cohere / Voyage / TEI / LLM"]
  R --> QA["AnswerWithContextPipeline + citations"]
  QA --> LLM["LLM (OpenAI, Gemini, Ollama, ...)"]
  G --> IDX["IndexDocumentPipeline: loaders + TokenSplitter"]
  IDX --> VS
  IDX --> DS
  G --> DB["SQLite: users, sources, settings"]
Component Path Role
Entry app.py, libs/ktem/ktem/main.py Builds the Gradio app and launches it with a queue
Settings flowsettings.py Stores, models, rerankers, reasoning pipelines, index types, feature flags
File index libs/ktem/ktem/index/file/ FileIndex, indexing and retrieval pipelines, per-index SQL tables
Graph indexes libs/ktem/ktem/index/file/graph/ Microsoft GraphRAG, nano-GraphRAG and LightRAG collections
Reasoning libs/ktem/ktem/reasoning/ FullQAPipeline, FullDecomposeQAPipeline, ReactAgentPipeline, RewooAgentPipeline
Retrieval core libs/kotaemon/kotaemon/indices/vectorindex.py VectorIndexing and VectorRetrieval
QA and citations libs/kotaemon/kotaemon/indices/qa/ Prompting, streaming, highlight and inline citations
Loaders libs/kotaemon/kotaemon/loaders/ PDF, Office, HTML, Unstructured, Adobe, Azure DI, Docling, PaddleOCR, Mathpix
Stores libs/kotaemon/kotaemon/storages/ Chroma, LanceDB, Qdrant, Milvus vector stores; LanceDB, Elasticsearch, file doc stores
Chat UI libs/ktem/ktem/pages/chat/ File selection, settings, streaming output channels

How a request flows

A question in the chat tab with the default “simple” reasoning:

  1. Build the pipeline. ChatPage.chat_fn calls create_pipeline. That asks every index for get_retriever_pipelines over the files selected in the UI, or uses the web-search retriever for the web command. It then instantiates the chosen reasoning class with those retrievers (chat/init.py). Output arrives as Documents on named channels (chat, info, plot, debug) and is rendered as it streams.
  2. Retrieve. FullQAPipeline.stream optionally rewrites the question, then runs each retriever on the raw message and dedupes by doc_id (simple.py). History-aware query contextualization is present but commented out, so follow-up questions are searched as typed.
  3. Scope. DocumentRetrievalPipeline.run returns nothing if no file is selected. Otherwise it maps the selected files to chunk ids through the index’s SQL Index table, adds a file_id IN (...) metadata filter and optional MMR, and calls VectorRetrieval with top_k 5 (pipelines.py).
  4. Hybrid search and rerank. In hybrid mode, VectorRetrieval.run queries the vector store and the doc store’s full-text index in two threads, fetching 10x top_k from each. It concatenates the keyword hits (score -1) ahead of the vector hits, runs each configured reranker in order, truncates to top_k, and attaches page thumbnails (vectorindex.py).
  5. Answer. PrepareEvidencePipeline packs the chunks into text, table or figure evidence. AnswerWithContextPipeline.stream starts the citation and mind-map LLM calls in background threads, builds system, history and prompt messages (with image URLs for figures in multimodal mode), and streams tokens from llm.stream, falling back to a blocking call. A qa_score is computed from logprobs when the model returns them (citation_qa.py).
  6. Show evidence. In parallel, the first retriever’s LLM scorer grades each chunk’s relevance. After the answer, show_citations_and_addons matches the extracted quotes back to the chunks for highlighting and renders the evidence panel (simple.py).

Key components

Indexing

IndexDocumentPipeline.route picks a loader by extension, or the web reader for URLs, and falls back to Unstructured for unknown types. It wraps the loader in an IndexPipeline with a TokenSplitter of 1,024 tokens and 256 overlap, splitting on blank lines first (pipelines.py). The per-user “File loader” setting swaps the PDF reader for Adobe, Azure Document Intelligence, Docling, PaddleOCR PP-StructureV3 or PaddleOCR-VL. The default PDF reader also renders page thumbnails (files.py). handle_docs splits only text documents. Tables, figures and thumbnails are indexed whole, and text chunks are linked to their page thumbnail. Embedding can run in a background thread (“quick index mode”) (pipelines.py).

Stores and models

The defaults are a Chroma vector store and a LanceDB document store (full-text search), both on local disk, plus SQLite at ktem_app_data/user_data/sql.db (flowsettings.py). LLM and embedding entries are generated from environment variables (OpenAI, Azure, Voyage, Ollama via LOCAL_MODEL, FastEmbed), with Claude, Gemini, Groq, Cohere and Mistral as placeholders to fill in through the UI’s resources tab.

Rerankers

Cohere rerank-v4.0-fast is the default reranker, and Voyage, TEI and LLM-based rankers are alternatives (flowsettings.py). The Cohere reranker silently returns the input unchanged when no API key is set (cohere.py). Combined with step 4, a keyless install ranks keyword hits first and cuts at five. Setting a reranker key matters more than the README suggests.

Citations

There are two modes. Highlight citation runs a separate function-calling CitationPipeline that extracts verbatim quotes, which are fuzzy-matched back to chunk spans for PDF highlighting. Inline citation (AnswerWithInlineCitation) makes the model emit numbered markers and start/end phrases in the answer itself.

Graph indexes and agents

GraphRAG, nano-GraphRAG and LightRAG appear as extra collection types, toggled by USE_* flags. The Microsoft GraphRAG index writes the documents to disk and shells out to python -m graphrag.index (graph/pipelines.py). FullDecomposeQAPipeline splits a question into sub-questions. The ReAct and ReWOO pipelines get tools such as SearchDoc, Wikipedia, Google and LLM, plus any registered MCP server’s tools (react.py).

Extending it

  • Settings first. Swap stores, models and rerankers by editing the __type__ class paths in flowsettings.py. Add or remove reasoning pipelines in KH_REASONINGS and index types in KH_INDEX_TYPES/KH_INDICES.
  • Plugins. The ktem_declare_extensions pluggy hook lets a package contribute reasoning pipelines or index types with their own callbacks and settings (extension_protocol.py).
  • New retrieval behaviour. Subclass IndexDocumentPipeline.route for custom loaders and splitters, or BaseFileIndexRetriever for a new retriever. Components declare settings that the UI renders automatically.
  • Agent tools. Register MCP servers in the UI, or add a tool to the agent tool registry.

Running it

  • Docker. The image has four targets: lite (slim Python), full (adds OCR, LibreOffice, Unstructured and torch), paddle and ollama. All app data lives under ktem_app_data/.
  • Local. Install libs/kotaemon and libs/ktem, set provider keys in .env, and run python app.py. It launches the Gradio app with a request queue and opens a browser (app.py). For fully local use, point LOCAL_MODEL at Ollama and pick FastEmbed or Ollama embeddings.
  • Accounts. User management is on by default with an admin/admin account unless the KH_FEATURE_USER_MANAGEMENT_* variables are set. SSO is behind KH_SSO_ENABLED.

Strengths and caveats

  • Strength: a usable product out of the box. Login, private and shared collections, a file manager, a PDF viewer with highlighted evidence, relevance scores and mind maps. Few open-source RAG projects ship this much UI.
  • Strength: parser choice. Six PDF loaders, including Docling and two PaddleOCR modes, are selectable per user without code.
  • Strength: experimentation surface. Graph RAG variants, decomposition, ReAct/ReWOO and MCP tools can be compared side by side in one app.
  • Caveat: naive hybrid merge. Keyword and vector results are concatenated, not fused. The in-function dedupe compares Document objects to id strings, so it never fires. Ranking quality depends on a configured reranker.
  • Caveat: UI-bound, single-process. No API, SQLite plus embedded stores, and background threads instead of a task queue. It scales to a team, not to a service.
  • Caveat: conversation context. Retrieval uses the literal last message unless the rewrite option is on, so short follow-ups like “and the second one?” retrieve poorly.
  • Caveat: weak defaults to change. A default admin/admin login and placeholder API keys in settings.

Sources: code at 9ad3e4e, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the RAG engines questions

Each answer was drafted by a code-reading agent at commit 9ad3e4e. Its citations were checked mechanically. Compare with the other rag engines →

How are documents parsed and chunked?

answered

Documents are parsed by file-type-specific readers (extractors) defined in KH_DEFAULT_FILE_EXTRACTORS (libs/kotaemon/kotaemon/indices/ingests/files.py, line 48–64). Supported formats: .pdf, .xlsx, .docx, .pptx, .xls, .doc, .html, .mhtml, .png, .jpeg, .jpg, .tiff, .tif, .txt, .md. The PDF reader can be swapped via reader_mode to use Adobe PDF Extract API, Azure AI Document Intelligence, or PaddleOCR modes (libs/ktem/ktem/index/file/pipelines.py, lines 677–706). In paddle-struct mode, the system uses PPStructureV3 for layout analysis, table extraction, and figure detection; paddle-vl uses a VLM for vision-language document parsing (libs/kotaemon/kotaemon/indices/ingests/files.py, lines 43–45). Table extraction is supported via these specialized readers — tables are stored with metadata type "table" and their HTML originals. Thumbnails are extracted per page for PDFs. After loading, the IndexPipeline.handle_docs() method splits text documents using the configured TokenSplitter (default: chunk_size=1024 tokens, chunk_overlap=256, separator `

, backup_separators [ , ., ​]) (libs/ktem/ktem/index/file/pipelines.py, lines 784–789). A SentenceWindowSplitter is also available (libs/kotaemon/kotaemon/indices/splitters/init.py`, lines 31–49). Non-text chunks (tables, images, thumbnails) bypass the splitter and are indexed alongside text chunks. Chunks are batched (200 at a time) into the docstore and vectorstore.

How are embeddings and indexes built and stored?

answered

Embeddings are produced by pluggable models: OpenAI, Azure OpenAI, Cohere, VoyageAI, Google, Mistral, HuggingFace (via LangChain wrappers), TEI (Text-Embedding-Inference) endpoints, and local FastEmbed models (e.g. BAAI/bge-small-en-v1.5) (libs/kotaemon/kotaemon/embeddings/ directory). The VectorIndexing class calls self.embedding(docs) to generate embeddings, then adds them to both a vector store and a doc store (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 84–93). Vector stores available: Chroma (persistent client), LanceDB, Qdrant, Milvus (with lazy init by dimension), and SimpleFileVectorStore (libs/kotaemon/kotaemon/storages/vectorstores/). Default in flowsettings.py is Chroma at ktem_app_data/user_data/vectorstore (line 100–103). Document stores support full-text search (BM25) via Elasticsearch (with custom BM25 similarity, line 38–46 of elasticsearch.py) or LanceDB (with FTS index using en_stem tokenizer, line 56–60 of lancedb.py). The default docstore in flowsettings.py is LanceDBDocumentStore (line 95–97). Hybrid retrieval is supported by pairing a vector store with a BM25-capable doc store. Metadata stored includes file_id, page_label, file_name, type, thumbnail_doc_id, and window for sentence-window splitters. The indexing also writes chunk content to a markdown cache directory and records source-to-chunk relationships in an SQL Index table (libs/ktem/ktem/index/file/pipelines.py, lines 426–464).

How is retrieval performed?

answered

Retrieval is performed by VectorRetrieval.run() (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 134–303), which supports three modes via retrieval_mode:

  • vector: embeds the query, queries the vector store, fetches full text from the doc store.
  • text: performs BM25 full-text search on the doc store (LanceDB FTS or Elasticsearch BM25).
  • hybrid: runs both in parallel threads, de-duplicates results (giving priority to vector hits with scores), and merges them (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 186–237). This is the default mode.

Results are re-ranked by configurable rerankers: CohereReranking (via Cohere's rerank API), LLMReranking (LLM-based YES/NO relevance filter), LLMTrulensScoring (0–10 relevance grader that normalizes to 0–1 and stores as llm_trulens_score), and VoyageAIReranking (libs/kotaemon/kotaemon/indices/rankings/). The reranker list is applied in order; if a reranker is LLMReranking, results are first truncated to top_k (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 240–245).

Filters are supported via LlamaIndex MetadataFilters with file_id in selected doc IDs (libs/ktem/ktem/index/file/pipelines.py, lines 158–167). MMR (Maximum Marginal Relevance) can be enabled for diversity (libs/ktem/ktem/index/file/pipelines.py, lines 169–172).

Query rewriting is optional: AddQueryContextPipeline uses the LLM to generate a search query from conversation history (libs/ktem/ktem/reasoning/simple.py, lines 42–83). Question decomposition is available via FullDecomposeQAPipeline which splits complex questions into sub-questions and retrieves per sub-question (libs/ktem/ktem/reasoning/simple.py, lines 487–609). Agentic retrieval is also supported: ReactAgentPipeline integrates a ReAct agent with tools including doc search, Wikipedia, Google, and LLM (libs/ktem/ktem/reasoning/react.py, lines 181–306). Web search retrievers (Tavily, Jina) are available (libs/kotaemon/kotaemon/indices/retrievers/).

Editor's note. Correction: hybrid mode does not fuse or de-duplicate; full-text hits (score -1) are concatenated ahead of vector hits and the in-function dedupe check never matches, so without a reranker key the top 5 are mostly keyword hits. History-based query contextualization (AddQueryContextPipeline) is commented out in FullQAPipeline.retrieve.

How are answers generated and grounded?

answered

Answers are generated via AnswerWithContextPipeline (or AnswerWithInlineCitation) which assembles the retrieved evidence and user question into a prompt (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 40–80). Four prompt templates exist: text, table, figure, and chatbot. The chosen template is populated with {context}, {question}, and {lang} variables. The conversation history (up to n_last_interactions) is included as message turn pairs (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 234–240). The system prompt is configurable via UI.

Streaming is the default path: the LLM's .stream() method yields Document(channel="chat") tokens in real-time. If streaming is unsupported, it falls back to .invoke() (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 262–272).

Citation/attribution works two ways:

  • Highlight citation: An LLM function-calling pipeline (CitationPipeline) extracts direct quotes from the evidence using CiteEvidence schema, then match_evidence_with_context() uses SequenceMatcher to find spans in original docs for PDF viewer highlighting (libs/kotaemon/kotaemon/indices/qa/citation.py, lines 22–95).
  • Inline citation: AnswerWithInlineCitation prompts the LLM to output 【N】 markers in the answer alongside START_PHRASE/END_PHRASE delimiting exact spans, then parses and links them to the source documents (libs/kotaemon/kotaemon/indices/qa/citation_qa_inline.py, lines 70–361).

Multimodal QA is supported: when evidence_mode is figure and use_multimodal is true, image URLs are injected into the HumanMessage (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 242–256). An optional LLMTrulensScoring relevance score is computed per document in a background thread (libs/ktem/ktem/reasoning/simple.py, lines 297–306). Mindmap generation and citation embedding visualizations run as optional background threads.

Agentic/multi-step answering is implemented via ReAct (libs/ktem/ktem/reasoning/react.py) and ReWOO (libs/ktem/ktem/reasoning/rewoo/) agents. The ReAct agent uses the Think-Action-Observation loop with tools (doc search, Wikipedia, Google, LLM, MCP tools).

How is quality evaluated or observed?

insufficient evidence

There is no built-in evaluation framework (no RAGAS, TruLens, DeepEval, or similar eval harness integrated). The system does provide runtime quality signals: (1) LLMTrulensScoring (libs/kotaemon/kotaemon/indices/rankings/llm_trulens.py) scores retrieved documents on a 0–10 relevance scale per query; (2) qa_score is derived from the exponential average of LLM response log-probabilities (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 274–277); (3) a context-relevance warning threshold (CONTEXT_RELEVANT_WARNING_SCORE) triggers a UI warning when the max LLM relevance score is below 0.3 (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, line 36–38). There is no built-in logging/tracing system (e.g., LangSmith, Weights & Biases, MLflow) connected to the pipeline — observability would need to be added externally. I checked the Kotaemon library and the ktem app modules; no evaluation harness or metric tracker was found.

How is it deployed and operated?

answered

Kotaemon is a self-hosted web application, not a library or API service. It runs on Gradio (app.py, line 18–26) and is served as a full UI. The entry point is python app.py which launches ktem.main.App with .queue().launch(). A Gradio server listens on configurable GRADIO_SERVER_NAME and GRADIO_SERVER_PORT (default 7860). The UI includes Chat, File management (indices), Settings, Resources, and Help tabs (libs/ktem/ktem/main.py, lines 46–82).

Docker deployment is the recommended path. The Dockerfile (Dockerfile) provides four variants: lite (Python 3.11-slim with basic dependencies + pdf.js), full (adds Tesseract OCR, LibreOffice, Unstructured, torch), paddle (adds PaddleOCR GPU support with CUDA 11.3), and ollama (bundles Ollama with nomic-embed-text). Images are published on GHCR (ghcr.io/cinnamon/kotaemon).

Infrastructure requirements: SQLite database (for users, indices, model configs), a filestorage directory for uploaded files, and persistent directories for vectorstore and docstore — all under ./ktem_app_data by default. Multi-user support is built in (SQLAlchemy-backed user management with login page, private/public collections, admin default user) (flowsettings.py, lines 76–84). SSO is optionally supported via KH_SSO_ENABLED. The first-run setup wizard can be enabled.

Scaling and multi-tenancy: The app uses SQLite by default (single-server). It supports private (per-user) and public file collections via the private index config and user_id column in Source/Index tables (libs/ktem/ktem/index/file/index.py, lines 62–126). The index_manager manages multiple indices (File Collection, GraphRAG collections, etc.) with separate tables, vector stores, and doc stores per index (libs/ktem/ktem/index/file/index.py, lines 153–163). Embedding and LLM models are managed via per-user managers (SQL-persisted) with a pool pattern (libs/ktem/ktem/embeddings/manager.py, libs/ktem/ktem/llms/manager.py). Elasticsearch is listed as an alternative docstore for deployments that need to scale beyond SQLite's full-text search capabilities. The app also provides a settings.yaml config file and fly.toml for Fly.io deployment.