Cinnamon/kotaemon
Gradio app for chatting with documents: keyword plus vector retrieval, rerankers, highlighted citations, GraphRAG/LightRAG indexes and ReAct agents.
Overview
Kotaemon, from Cinnamon, is a self-hosted Gradio web app for asking questions about your own documents. You upload files into “collections”, choose which files to search, and chat. Answers stream in with highlighted evidence in a built-in PDF viewer, an optional mind map, and relevance scores per retrieved chunk. It is aimed at two groups: end users who want a local “chat with PDFs” tool with multi-user login, and developers who want a hackable RAG reference app.
The repository has two Python packages. libs/kotaemon is the component library: loaders, splitters, embeddings, LLM wrappers, vector and document stores, rerankers, QA pipelines and agents. They are built on Cinnamon’s theflow Node/Param composition model, and many wrap LangChain or LlamaIndex classes. libs/ktem is the application: Gradio pages, SQLModel tables in SQLite for users, sources and settings, index management, reasoning pipelines, an MCP server registry, and a pluggy extension protocol. Everything is wired in flowsettings.py, a Python settings module of class paths and dicts.
There is no REST API. The product is the UI, and every pipeline runs inside the Gradio process.
Architecture
flowchart LR
U["Browser"] --> G["Gradio app (ktem.main.App)"]
G --> CHAT["ChatPage.chat_fn"]
CHAT --> R["Reasoning: FullQA / Decompose / ReAct / ReWOO"]
R --> RET["DocumentRetrievalPipeline per index"]
RET --> VR["VectorRetrieval (vector + full-text)"]
VR --> VS["Vector store: Chroma (default)"]
VR --> DS["Doc store: LanceDB FTS (default)"]
VR --> RR["Rerankers: Cohere / Voyage / TEI / LLM"]
R --> QA["AnswerWithContextPipeline + citations"]
QA --> LLM["LLM (OpenAI, Gemini, Ollama, ...)"]
G --> IDX["IndexDocumentPipeline: loaders + TokenSplitter"]
IDX --> VS
IDX --> DS
G --> DB["SQLite: users, sources, settings"]
| Component | Path | Role |
|---|---|---|
| Entry | app.py, libs/ktem/ktem/main.py |
Builds the Gradio app and launches it with a queue |
| Settings | flowsettings.py |
Stores, models, rerankers, reasoning pipelines, index types, feature flags |
| File index | libs/ktem/ktem/index/file/ |
FileIndex, indexing and retrieval pipelines, per-index SQL tables |
| Graph indexes | libs/ktem/ktem/index/file/graph/ |
Microsoft GraphRAG, nano-GraphRAG and LightRAG collections |
| Reasoning | libs/ktem/ktem/reasoning/ |
FullQAPipeline, FullDecomposeQAPipeline, ReactAgentPipeline, RewooAgentPipeline |
| Retrieval core | libs/kotaemon/kotaemon/indices/vectorindex.py |
VectorIndexing and VectorRetrieval |
| QA and citations | libs/kotaemon/kotaemon/indices/qa/ |
Prompting, streaming, highlight and inline citations |
| Loaders | libs/kotaemon/kotaemon/loaders/ |
PDF, Office, HTML, Unstructured, Adobe, Azure DI, Docling, PaddleOCR, Mathpix |
| Stores | libs/kotaemon/kotaemon/storages/ |
Chroma, LanceDB, Qdrant, Milvus vector stores; LanceDB, Elasticsearch, file doc stores |
| Chat UI | libs/ktem/ktem/pages/chat/ |
File selection, settings, streaming output channels |
How a request flows
A question in the chat tab with the default “simple” reasoning:
- Build the pipeline.
ChatPage.chat_fncallscreate_pipeline. That asks every index forget_retriever_pipelinesover the files selected in the UI, or uses the web-search retriever for the web command. It then instantiates the chosen reasoning class with those retrievers (chat/init.py). Output arrives asDocuments on named channels (chat,info,plot,debug) and is rendered as it streams. - Retrieve.
FullQAPipeline.streamoptionally rewrites the question, then runs each retriever on the raw message and dedupes bydoc_id(simple.py). History-aware query contextualization is present but commented out, so follow-up questions are searched as typed. - Scope.
DocumentRetrievalPipeline.runreturns nothing if no file is selected. Otherwise it maps the selected files to chunk ids through the index’s SQLIndextable, adds afile_id IN (...)metadata filter and optional MMR, and callsVectorRetrievalwithtop_k5 (pipelines.py). - Hybrid search and rerank. In
hybridmode,VectorRetrieval.runqueries the vector store and the doc store’s full-text index in two threads, fetching 10xtop_kfrom each. It concatenates the keyword hits (score -1) ahead of the vector hits, runs each configured reranker in order, truncates totop_k, and attaches page thumbnails (vectorindex.py). - Answer.
PrepareEvidencePipelinepacks the chunks into text, table or figure evidence.AnswerWithContextPipeline.streamstarts the citation and mind-map LLM calls in background threads, builds system, history and prompt messages (with image URLs for figures in multimodal mode), and streams tokens fromllm.stream, falling back to a blocking call. Aqa_scoreis computed from logprobs when the model returns them (citation_qa.py). - Show evidence. In parallel, the first retriever’s LLM scorer grades each chunk’s relevance. After the answer,
show_citations_and_addonsmatches the extracted quotes back to the chunks for highlighting and renders the evidence panel (simple.py).
Key components
Indexing
IndexDocumentPipeline.route picks a loader by extension, or the web reader for URLs, and falls back to Unstructured for unknown types. It wraps the loader in an IndexPipeline with a TokenSplitter of 1,024 tokens and 256 overlap, splitting on blank lines first (pipelines.py). The per-user “File loader” setting swaps the PDF reader for Adobe, Azure Document Intelligence, Docling, PaddleOCR PP-StructureV3 or PaddleOCR-VL. The default PDF reader also renders page thumbnails (files.py). handle_docs splits only text documents. Tables, figures and thumbnails are indexed whole, and text chunks are linked to their page thumbnail. Embedding can run in a background thread (“quick index mode”) (pipelines.py).
Stores and models
The defaults are a Chroma vector store and a LanceDB document store (full-text search), both on local disk, plus SQLite at ktem_app_data/user_data/sql.db (flowsettings.py). LLM and embedding entries are generated from environment variables (OpenAI, Azure, Voyage, Ollama via LOCAL_MODEL, FastEmbed), with Claude, Gemini, Groq, Cohere and Mistral as placeholders to fill in through the UI’s resources tab.
Rerankers
Cohere rerank-v4.0-fast is the default reranker, and Voyage, TEI and LLM-based rankers are alternatives (flowsettings.py). The Cohere reranker silently returns the input unchanged when no API key is set (cohere.py). Combined with step 4, a keyless install ranks keyword hits first and cuts at five. Setting a reranker key matters more than the README suggests.
Citations
There are two modes. Highlight citation runs a separate function-calling CitationPipeline that extracts verbatim quotes, which are fuzzy-matched back to chunk spans for PDF highlighting. Inline citation (AnswerWithInlineCitation) makes the model emit numbered markers and start/end phrases in the answer itself.
Graph indexes and agents
GraphRAG, nano-GraphRAG and LightRAG appear as extra collection types, toggled by USE_* flags. The Microsoft GraphRAG index writes the documents to disk and shells out to python -m graphrag.index (graph/pipelines.py). FullDecomposeQAPipeline splits a question into sub-questions. The ReAct and ReWOO pipelines get tools such as SearchDoc, Wikipedia, Google and LLM, plus any registered MCP server’s tools (react.py).
Extending it
- Settings first. Swap stores, models and rerankers by editing the
__type__class paths inflowsettings.py. Add or remove reasoning pipelines inKH_REASONINGSand index types inKH_INDEX_TYPES/KH_INDICES. - Plugins. The
ktem_declare_extensionspluggy hook lets a package contribute reasoning pipelines or index types with their own callbacks and settings (extension_protocol.py). - New retrieval behaviour. Subclass
IndexDocumentPipeline.routefor custom loaders and splitters, orBaseFileIndexRetrieverfor a new retriever. Components declare settings that the UI renders automatically. - Agent tools. Register MCP servers in the UI, or add a tool to the agent tool registry.
Running it
- Docker. The image has four targets:
lite(slim Python),full(adds OCR, LibreOffice, Unstructured and torch),paddleandollama. All app data lives underktem_app_data/. - Local. Install
libs/kotaemonandlibs/ktem, set provider keys in.env, and runpython app.py. It launches the Gradio app with a request queue and opens a browser (app.py). For fully local use, pointLOCAL_MODELat Ollama and pick FastEmbed or Ollama embeddings. - Accounts. User management is on by default with an
admin/adminaccount unless theKH_FEATURE_USER_MANAGEMENT_*variables are set. SSO is behindKH_SSO_ENABLED.
Strengths and caveats
- Strength: a usable product out of the box. Login, private and shared collections, a file manager, a PDF viewer with highlighted evidence, relevance scores and mind maps. Few open-source RAG projects ship this much UI.
- Strength: parser choice. Six PDF loaders, including Docling and two PaddleOCR modes, are selectable per user without code.
- Strength: experimentation surface. Graph RAG variants, decomposition, ReAct/ReWOO and MCP tools can be compared side by side in one app.
- Caveat: naive hybrid merge. Keyword and vector results are concatenated, not fused. The in-function dedupe compares
Documentobjects to id strings, so it never fires. Ranking quality depends on a configured reranker. - Caveat: UI-bound, single-process. No API, SQLite plus embedded stores, and background threads instead of a task queue. It scales to a team, not to a service.
- Caveat: conversation context. Retrieval uses the literal last message unless the rewrite option is on, so short follow-ups like “and the second one?” retrieve poorly.
- Caveat: weak defaults to change. A default
admin/adminlogin and placeholder API keys in settings.
Sources: code at 9ad3e4e, deepwiki-open wiki (12 pages), verified Q&A.
How it answers the RAG engines questions
Each answer was drafted by a code-reading agent at commit 9ad3e4e. Its citations were checked mechanically. Compare with the other rag engines →
How are documents parsed and chunked?
answeredDocuments are parsed by file-type-specific readers (extractors) defined in KH_DEFAULT_FILE_EXTRACTORS (libs/kotaemon/kotaemon/indices/ingests/files.py, line 48–64). Supported formats: .pdf, .xlsx, .docx, .pptx, .xls, .doc, .html, .mhtml, .png, .jpeg, .jpg, .tiff, .tif, .txt, .md. The PDF reader can be swapped via reader_mode to use Adobe PDF Extract API, Azure AI Document Intelligence, or PaddleOCR modes (libs/ktem/ktem/index/file/pipelines.py, lines 677–706). In paddle-struct mode, the system uses PPStructureV3 for layout analysis, table extraction, and figure detection; paddle-vl uses a VLM for vision-language document parsing (libs/kotaemon/kotaemon/indices/ingests/files.py, lines 43–45). Table extraction is supported via these specialized readers — tables are stored with metadata type "table" and their HTML originals. Thumbnails are extracted per page for PDFs. After loading, the IndexPipeline.handle_docs() method splits text documents using the configured TokenSplitter (default: chunk_size=1024 tokens, chunk_overlap=256, separator `
, backup_separators [
, ., ]) (libs/ktem/ktem/index/file/pipelines.py, lines 784–789). A SentenceWindowSplitter is also available (libs/kotaemon/kotaemon/indices/splitters/init.py`, lines 31–49). Non-text chunks (tables, images, thumbnails) bypass the splitter and are indexed alongside text chunks. Chunks are batched (200 at a time) into the docstore and vectorstore.
How are embeddings and indexes built and stored?
answeredEmbeddings are produced by pluggable models: OpenAI, Azure OpenAI, Cohere, VoyageAI, Google, Mistral, HuggingFace (via LangChain wrappers), TEI (Text-Embedding-Inference) endpoints, and local FastEmbed models (e.g. BAAI/bge-small-en-v1.5) (libs/kotaemon/kotaemon/embeddings/ directory). The VectorIndexing class calls self.embedding(docs) to generate embeddings, then adds them to both a vector store and a doc store (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 84–93). Vector stores available: Chroma (persistent client), LanceDB, Qdrant, Milvus (with lazy init by dimension), and SimpleFileVectorStore (libs/kotaemon/kotaemon/storages/vectorstores/). Default in flowsettings.py is Chroma at ktem_app_data/user_data/vectorstore (line 100–103). Document stores support full-text search (BM25) via Elasticsearch (with custom BM25 similarity, line 38–46 of elasticsearch.py) or LanceDB (with FTS index using en_stem tokenizer, line 56–60 of lancedb.py). The default docstore in flowsettings.py is LanceDBDocumentStore (line 95–97). Hybrid retrieval is supported by pairing a vector store with a BM25-capable doc store. Metadata stored includes file_id, page_label, file_name, type, thumbnail_doc_id, and window for sentence-window splitters. The indexing also writes chunk content to a markdown cache directory and records source-to-chunk relationships in an SQL Index table (libs/ktem/ktem/index/file/pipelines.py, lines 426–464).
How is retrieval performed?
answeredRetrieval is performed by VectorRetrieval.run() (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 134–303), which supports three modes via retrieval_mode:
- vector: embeds the query, queries the vector store, fetches full text from the doc store.
- text: performs BM25 full-text search on the doc store (LanceDB FTS or Elasticsearch BM25).
- hybrid: runs both in parallel threads, de-duplicates results (giving priority to vector hits with scores), and merges them (
libs/kotaemon/kotaemon/indices/vectorindex.py, lines 186–237). This is the default mode.
Results are re-ranked by configurable rerankers: CohereReranking (via Cohere's rerank API), LLMReranking (LLM-based YES/NO relevance filter), LLMTrulensScoring (0–10 relevance grader that normalizes to 0–1 and stores as llm_trulens_score), and VoyageAIReranking (libs/kotaemon/kotaemon/indices/rankings/). The reranker list is applied in order; if a reranker is LLMReranking, results are first truncated to top_k (libs/kotaemon/kotaemon/indices/vectorindex.py, lines 240–245).
Filters are supported via LlamaIndex MetadataFilters with file_id in selected doc IDs (libs/ktem/ktem/index/file/pipelines.py, lines 158–167). MMR (Maximum Marginal Relevance) can be enabled for diversity (libs/ktem/ktem/index/file/pipelines.py, lines 169–172).
Query rewriting is optional: AddQueryContextPipeline uses the LLM to generate a search query from conversation history (libs/ktem/ktem/reasoning/simple.py, lines 42–83). Question decomposition is available via FullDecomposeQAPipeline which splits complex questions into sub-questions and retrieves per sub-question (libs/ktem/ktem/reasoning/simple.py, lines 487–609). Agentic retrieval is also supported: ReactAgentPipeline integrates a ReAct agent with tools including doc search, Wikipedia, Google, and LLM (libs/ktem/ktem/reasoning/react.py, lines 181–306). Web search retrievers (Tavily, Jina) are available (libs/kotaemon/kotaemon/indices/retrievers/).
How are answers generated and grounded?
answeredAnswers are generated via AnswerWithContextPipeline (or AnswerWithInlineCitation) which assembles the retrieved evidence and user question into a prompt (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 40–80). Four prompt templates exist: text, table, figure, and chatbot. The chosen template is populated with {context}, {question}, and {lang} variables. The conversation history (up to n_last_interactions) is included as message turn pairs (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 234–240). The system prompt is configurable via UI.
Streaming is the default path: the LLM's .stream() method yields Document(channel="chat") tokens in real-time. If streaming is unsupported, it falls back to .invoke() (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 262–272).
Citation/attribution works two ways:
- Highlight citation: An LLM function-calling pipeline (
CitationPipeline) extracts direct quotes from the evidence usingCiteEvidenceschema, thenmatch_evidence_with_context()uses SequenceMatcher to find spans in original docs for PDF viewer highlighting (libs/kotaemon/kotaemon/indices/qa/citation.py, lines 22–95). - Inline citation:
AnswerWithInlineCitationprompts the LLM to output【N】markers in the answer alongside START_PHRASE/END_PHRASE delimiting exact spans, then parses and links them to the source documents (libs/kotaemon/kotaemon/indices/qa/citation_qa_inline.py, lines 70–361).
Multimodal QA is supported: when evidence_mode is figure and use_multimodal is true, image URLs are injected into the HumanMessage (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 242–256). An optional LLMTrulensScoring relevance score is computed per document in a background thread (libs/ktem/ktem/reasoning/simple.py, lines 297–306). Mindmap generation and citation embedding visualizations run as optional background threads.
Agentic/multi-step answering is implemented via ReAct (libs/ktem/ktem/reasoning/react.py) and ReWOO (libs/ktem/ktem/reasoning/rewoo/) agents. The ReAct agent uses the Think-Action-Observation loop with tools (doc search, Wikipedia, Google, LLM, MCP tools).
How is quality evaluated or observed?
insufficient evidenceThere is no built-in evaluation framework (no RAGAS, TruLens, DeepEval, or similar eval harness integrated). The system does provide runtime quality signals: (1) LLMTrulensScoring (libs/kotaemon/kotaemon/indices/rankings/llm_trulens.py) scores retrieved documents on a 0–10 relevance scale per query; (2) qa_score is derived from the exponential average of LLM response log-probabilities (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, lines 274–277); (3) a context-relevance warning threshold (CONTEXT_RELEVANT_WARNING_SCORE) triggers a UI warning when the max LLM relevance score is below 0.3 (libs/kotaemon/kotaemon/indices/qa/citation_qa.py, line 36–38). There is no built-in logging/tracing system (e.g., LangSmith, Weights & Biases, MLflow) connected to the pipeline — observability would need to be added externally. I checked the Kotaemon library and the ktem app modules; no evaluation harness or metric tracker was found.
How is it deployed and operated?
answeredKotaemon is a self-hosted web application, not a library or API service. It runs on Gradio (app.py, line 18–26) and is served as a full UI. The entry point is python app.py which launches ktem.main.App with .queue().launch(). A Gradio server listens on configurable GRADIO_SERVER_NAME and GRADIO_SERVER_PORT (default 7860). The UI includes Chat, File management (indices), Settings, Resources, and Help tabs (libs/ktem/ktem/main.py, lines 46–82).
Docker deployment is the recommended path. The Dockerfile (Dockerfile) provides four variants: lite (Python 3.11-slim with basic dependencies + pdf.js), full (adds Tesseract OCR, LibreOffice, Unstructured, torch), paddle (adds PaddleOCR GPU support with CUDA 11.3), and ollama (bundles Ollama with nomic-embed-text). Images are published on GHCR (ghcr.io/cinnamon/kotaemon).
Infrastructure requirements: SQLite database (for users, indices, model configs), a filestorage directory for uploaded files, and persistent directories for vectorstore and docstore — all under ./ktem_app_data by default. Multi-user support is built in (SQLAlchemy-backed user management with login page, private/public collections, admin default user) (flowsettings.py, lines 76–84). SSO is optionally supported via KH_SSO_ENABLED. The first-run setup wizard can be enabled.
Scaling and multi-tenancy: The app uses SQLite by default (single-server). It supports private (per-user) and public file collections via the private index config and user_id column in Source/Index tables (libs/ktem/ktem/index/file/index.py, lines 62–126). The index_manager manages multiple indices (File Collection, GraphRAG collections, etc.) with separate tables, vector stores, and doc stores per index (libs/ktem/ktem/index/file/index.py, lines 153–163). Embedding and LLM models are managed via per-user managers (SQL-persisted) with a pool pattern (libs/ktem/ktem/embeddings/manager.py, libs/ktem/ktem/llms/manager.py). Elasticsearch is listed as an alternative docstore for deployments that need to scale beyond SQLite's full-text search capabilities. The app also provides a settings.yaml config file and fly.toml for Fly.io deployment.