LLMs Technical Reviews
Home / AI web scraping / omniparse

adithya-s-k/omniparse

Self-hosted FastAPI server that wraps marker, Florence-2, Whisper and a Selenium crawler to turn files and URLs into markdown.

GitHub ↗★ 7.9kPythonGPL-3.0commit 9d1ae83 · 2024-11-04homepage ↗

Overview

OmniParse is a self-hosted FastAPI server that turns uploaded files and web pages into markdown plus base64 images for LLM and RAG pipelines. It does very little parsing of its own. It is glue around existing tools: marker-pdf (Surya OCR, layout and reading order) for PDFs, LibreOffice for Office files, Microsoft’s Florence-2 for vision tasks, OpenAI Whisper for audio and video, and a Selenium plus BeautifulSoup crawler adapted from an early crawl4ai for web pages. Every endpoint returns the same responseDocument shape: text, images, metadata and chunks.

The pinned commit is from November 2024, and the code reads like a mid-2024 prototype that stopped there. The document and media paths work as thin wrappers. The web path renders one URL with headless Chrome and cleans it with heuristics. The LLM block extraction, the chunking strategies, the /crawl and /search routes and the cache flag exist in the tree but are not wired to any request. For an “ai-scraping” tool, the main thing to know is that the shipped server makes no LLM calls.

It suits someone who wants one local HTTP endpoint for “PDF, DOCX, image, MP3 or URL in, markdown out” and has a CUDA GPU. It is not a crawler and not an LLM extraction framework.

Architecture

flowchart LR
  C["Client (curl, SDK, Gradio)"] --> S["server.py FastAPI app"]
  S --> D["/parse_document router"]
  S --> I["/parse_image router"]
  S --> M["/parse_media router"]
  S --> W["/parse_website router"]
  S --> G["Gradio demo_ui at /"]
  D --> LO["LibreOffice to PDF"]
  LO --> MK["marker convert_single_pdf"]
  D --> MK
  I --> IP["img2pdf"] --> MK
  I --> FL["Florence-2 tasks"]
  M --> WH["Whisper transcribe"]
  W --> SE["Selenium Chrome"]
  SE --> CL["BeautifulSoup clean + html2text"]
  MK --> R["responseDocument JSON"]
  FL --> R
  WH --> R
  CL --> R
Component Path Role
Server entry server.py Builds the FastAPI app, CORS, mounts four routers and the Gradio UI, parses CLI flags, starts uvicorn
Model state omniparse/__init__.py SharedState singleton and load_omnimodel that loads marker models, Florence-2, Whisper and the web crawler
Response model omniparse/models/__init__.py responseDocument and responseImage Pydantic models
Documents omniparse/documents/router.py /pdf, /ppt, /docs and catch-all endpoints; LibreOffice conversion then marker
Images omniparse/image/ parse_image (image to PDF to marker) and process_image (Florence-2 task prompts)
Media omniparse/media/ Whisper transcription; video audio extracted with moviepy
Web omniparse/web/ WebCrawler, LocalSeleniumCrawlerStrategy, HTML cleaning, unused LLM block extraction
Gradio UI omniparse/demo.py Tabs that POST back to the same server
Python SDK python-sdk/omniparse_client/ Sync and async HTTP clients and a ParsedDocument saver

How a request flows

Take POST /parse_website/parse?url=https://example.com on a server started with --web:

  1. Startup. main() parses --documents, --media and --web, calls load_omnimodel, and starts uvicorn on server:app (server.py). With --web, load_omnimodel creates one WebCrawler, which starts one Chrome session through webdriver_manager (init.py, crawler_strategy.py).
  2. Route. parse_website takes url as a query parameter and awaits parse_url (router.py).
  3. Off the event loop. parse_url hardcodes word_count_threshold=5, screenshot=True and no CSS selector, then runs the synchronous WebCrawler.run in a ThreadPoolExecutor (web/init.py).
  4. Render. crawl() calls driver.get(url), waits up to 10 s for the html element, optionally runs injected JS, and returns page_source. take_screenshot() resizes the window to the full scroll size and returns a base64 JPEG (crawler_strategy.py).
  5. Clean. process_html calls get_content_of_website and extract_metadata (web_crawler.py). The cleaner collects links and media, strips scripts and attributes, swaps images for their alt text, prunes elements under five words, flattens nesting and converts to markdown with a CustomHTML2Text that fences <pre> blocks (utils.py).
  6. Respond. run() puts the markdown in text, the whole processing dict (raw HTML, cleaned HTML, links, media, metadata) in metadata, and the screenshot in images (web_crawler.py).

A document upload is shorter: /parse_document checks the extension, converts PPT/DOC/DOCX with libreoffice --headless --convert-to pdf, and calls marker’s convert_single_pdf with the preloaded models (documents/router.py).

Key components

Document and image parsing

All document quality comes from marker-pdf. OmniParse passes it bytes or a path and repackages full_text, images and out_meta. encode_images writes each extracted image to disk as PNG, reads it back as base64 and deletes it (utils.py). /parse_image/image reuses the same OCR path: it converts the image to a one-page PDF with img2pdf and sends that to marker (image/init.py).

Florence-2 vision tasks

/parse_image/process_image maps a task name such as Caption, Object Detection or OCR with Region to a Florence-2 prompt token and, for detection or segmentation tasks, draws boxes or polygons onto a copy of the image (process.py). Two details matter. The three “+ Grounding” options map to plain caption tokens, so they return captions without boxes. And run_example always moves inputs to "cuda", even though the model was loaded on whatever device was available (process.py).

Media transcription

Audio is written to a temp .wav and passed to Whisper small with fallback temperatures from 0.0 to 1.0 in steps of 0.2. Video first has its audio track written to MP3 with moviepy (media/init.py, media/utils.py). Only the transcript text is returned; segments and timestamps are dropped.

Dormant LLM extraction

omniparse/web/utils.py still carries crawl4ai’s LLM block extractor: extract_blocks fills PROMPT_EXTRACT_BLOCKS with the URL and sanitised HTML and calls litellm with up to three retries on rate limits, and process_sections fans sections out sequentially for Groq and in a thread pool otherwise (utils.py, L595-L612). The provider table lists OpenAI, Anthropic, Groq and Ollama models (config.py). Nothing in WebCrawler.run or any router calls these functions. The omniparse/chunking strategies are likewise never imported, so chunks is always empty.

Extending it

  • New input type. Add a router module, a parse function that returns responseDocument, a slot on SharedState for its model, and an include_router call in server.py.
  • Different web backend. CrawlerStrategy is a three-method ABC (crawl, take_screenshot, update_user_agent), and WebCrawler accepts any implementation (crawler_strategy.py, web_crawler.py).
  • Turn on LLM blocks. Call extract_blocks or process_sections on cleaned_html inside process_html and store the result in extracted_content, which today is always None.
  • Use it from Python. The repo has a python-sdk/ client, but check its routes first (see caveats). Calling the HTTP endpoints directly with requests or httpx is safer.

Running it

  • Docker. The Dockerfile is CUDA 11.8 on Ubuntu 22.04 with LibreOffice, ffmpeg, Chrome and ChromeDriver, and it starts python server.py --host 0.0.0.0 --port 8000 --documents --media --web (Dockerfile).
  • From source. Python 3.10+, Poetry dependencies including torch, marker-pdf ^0.2.16, surya-ocr, openai-whisper, selenium, gradio and flash-attn (pyproject.toml). download.py pre-fetches the models with the same flags.
  • Flags. --documents loads marker models and Florence-2 (also needed for /parse_image), --media loads Whisper, and --web starts Chrome. The routers are mounted unconditionally when the module loads; the flags only decide which models load and whether a router appears in the OpenAPI schema. Calling an endpoint whose model is not loaded fails at request time.

Strengths and caveats

  • Strength: one endpoint shape for many formats. PDF, Office, images, audio, video and URLs all come back as the same JSON, and everything runs locally with open models.
  • Strength: small and readable. The whole server is a few hundred lines of glue, easy to fork and rewire.
  • Caveat: no LLM in the live path. Block extraction, chunking, /crawl and /search (which return {"Coming soon"}) are scaffolding. The web output is heuristic cleaning only.
  • Caveat: GPU assumed. Florence-2 inference hardcodes CUDA, and the Docker image is CUDA-based. CPU runs are not supported for the vision tasks.
  • Caveat: one shared browser. A single Selenium driver serves every request, there is no cache despite the bypass_cache parameter, and there are no proxy, stealth or rate-limit features.
  • Caveat: SDK drift. The async client posts process_image to /parse_media/process_image and parse_website as a JSON body to /parse_website, neither of which matches the server, and parse_document never returns its result (omniparse.py, L401-L454).
  • Caveat: licensing and maintenance. The repo ships a GPL-3.0 LICENSE while pyproject.toml says “Apache”, and the last commit is from late 2024 with pins against 2024-era marker and transformers.

Sources: code at 9d1ae83, deepwiki-open wiki (12 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 9d1ae83. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Pages are fetched via Selenium ChromeDriver in headless mode, not plain HTTP. The LocalSeleniumCrawlerStrategy class (omniparse/web/crawler_strategy.py:64) creates a webdriver.Chrome instance with --headless, --no-sandbox, --disable-gpu flags. The crawl() method (line 104–139) calls driver.get(url) then waits up to 10 seconds for the html element to be present (EC.presence_of_all_elements_located). Optionally, custom JS code can be injected and awaited with a second WebDriverWait for document.readyState === "complete". A take_screenshot() method (line 141–198) captures the full page as a base64 JPEG. The WebCrawler.run() method (omniparse/web/web_crawler.py:112–151) drives the full pipeline: crawl HTML, then pass it to process_html() for extraction. PDF, PPT/DOCX, and images are not fetched via HTTP — they are uploaded as file bytes to the REST API; PPT/DOCX are converted to PDF via LibreOffice (omniparse/documents/router.py:67–72). There is no robots.txt check, no sitemap discovery, and no depth-based crawling in the current code.

How is content extracted or converted?

answered

HTML content is extracted via a multi-stage pipeline in get_content_of_website() (omniparse/web/utils.py:175–370). BeautifulSoup parses the HTML; the <body> is isolated and optionally narrowed with a CSS selector. Scripts, styles, and meta tags are removed. All HTML attributes are stripped except on <img> tags. Images are replaced with their alt text (or removed if missing). Empty tags, tags with fewer than MIN_WORD_THRESHOLD (5) words, and nested single-child elements are recursively pruned. The cleaned HTML is then converted to markdown via a custom CustomHTML2Text class (line 147–172, subclassing html2text.HTML2Text) that handles <pre> blocks with triple-backtick fences. Metadata (title, description, keywords, OG tags, Twitter cards) is extracted separately via BeautifulSoup in extract_metadata() (line 373–414). For documents, PDFs are processed by the marker-pdf library (convert_single_pdf), which runs Surya OCR for text detection, layout analysis, and reading-order reconstruction. PPT/DOCX files are first converted to PDF via LibreOffice (omniparse/documents/router.py:67), then run through the same marker pipeline. For images (omniparse/image/__init__.py:49–103), the image is converted to a PDF via img2pdf then parsed by marker-pdf's convert_single_pdf — effectively using the OCR pipeline. The SDK client's ParsedDocument (python-sdk/omniparse_client/utils.py:77) optionally extracts tables from the markdown output using regex (extract_markdown_tables, line 153). There is no schema-based extraction (the README lists it as a roadmap item, not implemented).

How are LLMs used, if at all?

answered

LLMs are used exclusively for the optional block-extraction feature in the web parser, gated behind the extract_blocks / extract_blocks_batch functions (omniparse/web/utils.py:473–560). These functions send cleaned HTML to an LLM via the litellm library (line 438–470), which acts as a unified interface to multiple providers. The supported providers are configured in omniparse/web/config.py:24–34 and include OpenAI (gpt-3.5-turbo, gpt-4-turbo, gpt-4o), Anthropic (claude-3-haiku, claude-3-opus, claude-3-sonnet), Groq (llama3-70b, llama3-8b), and Ollama (llama3). The default is openai/gpt-4-turbo (line 21). The system prompts (omniparse/web/prompts.py) ask the LLM to break the HTML into semantically relevant blocks, tagging each with an index, content, and (in the first variant) questions. For large pages, merge_chunks_based_on_token_threshold() (line 563–592) merges text chunks at a CHUNK_TOKEN_THRESHOLD of ~1000 tokens. Retries use exponential backoff (2s, 4s, 8s) on RateLimitError (line 444–470). Groq providers process sequentially with a 500ms delay; other providers use ThreadPoolExecutor parallelism (line 597–611). The extract_blocks_batch function (line 514–560) supports batched LLM calls via litellm.batch_completion. Notably, the core document/image/audio extraction pipeline does NOT use any LLM — it relies entirely on local models (Surya OCR, Florence-2, Whisper).

Editor's note. Correction: extract_blocks, extract_blocks_batch and process_sections are never called by WebCrawler.run or any router, so the shipped server makes no LLM calls at all; the litellm code is dormant crawl4ai scaffolding.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-bot measures are minimal. The web crawler (crawler_strategy.py:67–97) configures Chrome with headless=True, --no-sandbox, --disable-gpu, and --log-level=3 stealth flags but does not apply dedicated anti-detection patches like --disable-blink-features=AutomationControlled or spoof the Sec-CH-UA headers. The update_user_agent() method (line 99–102) allows setting a custom user-agent string, which is the only fingerprint-spoofing mechanism — it quits the existing driver and creates a new Chrome instance with the updated agent. There is no proxy rotation, no residential IP pool integration, no CAPTCHA solving, and no request rate-limiting logic in the code. The WebCrawler.run() method (web_crawler.py:112) has a bypass_cache parameter but this controls local caching, not server-side evasion. The only rate-conscious logic is in process_sections() (web_utils.py:597–601), which inserts a 500ms delay between Groq LLM API calls — this is API rate-limit compliance on the provider side, not anti-bot evasion.

How is crawling at scale implemented?

answered

Crawling at scale is not implemented in this codebase. The WebCrawler class (omniparse/web/web_crawler.py:32–151) is a single-page fetcher: run() takes one URL and returns one responseDocument. The fetch_pages() method (line 79–110) extends this to multiple URLs using Python's ThreadPoolExecutor to map the fetch_page call across a list of UrlModel objects — this is the only concurrency mechanism. There is no crawl queue (no Redis/Bull/Kafka), no URL deduplication set, no depth/scope limiting, no robots.txt parser, no politeness delay between requests, no sitemap discovery, and no distributed worker architecture. The /crawl and /search endpoints (web/router.py:24–31) both return {"Coming soon"} — these are placeholders. README roadmap items like "Batch processing data" and "Web crawl and search" are also unfulfilled. In short: the project can fetch a list of URLs concurrently via ThreadPoolExecutor but has none of the infrastructure needed for production-scale crawling.

What is the developer interface?

answered

OmniParse exposes four interfaces. (1) REST API — a FastAPI server (server.py) with four routers: /parse_document (PDF, PPT, DOCX via marker-pdf), /parse_image (image parsing and vision-task processing via Florence-2), /parse_media (audio/video transcription via Whisper), and /parse_website (URL parsing via Selenium). All endpoints return a responseDocument JSON (omniparse/models/__init__.py:15) containing text (markdown string), images (list of base64-encoded images), metadata (dict), and chunks (list of text chunks). (2) Python SDK client — python-sdk/omniparse_client/omniparse.py provides both a synchronous OmniParse class and an AsyncOmniParse class with methods like parse_document(), parse_pdf(), parse_image(), parse_website(), and process_image(). The async client uses httpx and aiofiles. The SDK outputs a ParsedDocument (utils.py:77) with optional file-saving to disk. (3) Gradio UI — omniparse/demo.py defines a multi-tab Gradio interface (Documents, Images/Parse/Process, Media, Web) that sends requests to the local FastAPI server and renders markdown, extracted images, and JSON output. (4) CLI — the server is launched via python server.py --host 0.0.0.0 --port 8000 --documents --media --web, with flags to selectively load models. There is no MCP server, no JavaScript bindings (the demo shows "Coming soon" for Python and JS code snippets), and no native TypeScript client. Output is always a responseDocument JSON with markdown text, optional images as base64, and metadata.

Editor's note. Correction: the async SDK is out of sync with the server: process_image posts to /parse_media/process_image (server route is /parse_image/process_image), parse_website sends a JSON body to /parse_website (server expects a url query parameter on /parse_website/parse), and parse_document never returns its result.