LLMs Technical Reviews
Home / AI web scraping / parsera

raznem/parsera

Python library that renders a page in headless Firefox, converts it to Markdown and has an LLM return the fields you describe as JSON.

GitHub ↗★ 1.4kPythonGPL-2.0commit 865d777 · 2025-12-17homepage ↗

Overview

Parsera is a small Python library for “describe the fields, get JSON back” scraping. You give it a URL and a dict such as {"Title": "News title", "Points": "Number of points"}. It opens the page in headless Firefox through Playwright, grabs the rendered HTML (iframes included), turns it into Markdown with markdownify, and asks a language model to return a list of records with those keys. There are no selectors, no XPath and no learned extraction rules. Every run re-reads the page with the model.

The whole package is about 1,100 lines of code. The interesting decision is where the LLM lives. A bare Parsera() does not call OpenAI at all. It sends the page to the vendor’s hosted endpoint, https://api.parsera.org/v1/parse, authenticated with PARSERA_API_KEY (parsera.py, api_extractor.py). Pass any LangChain chat model and extraction runs locally instead, through a chunking extractor with its own prompts. So the open-source library is both a client for the commercial Parsera service and a self-contained LLM scraper.

It suits one-off or low-volume jobs where the page layout changes often and per-page LLM cost is acceptable: a listing page, a product table, a news front page. It is a single-URL tool. There is no crawler, queue, link following or retry logic.

Architecture

flowchart LR
  U["Your code / CLI"] --> P["Parsera.run / arun"]
  P --> L["PageLoader"]
  L --> FF["Headless Firefox (Playwright)"]
  L --> HTML["HTML + iframe HTML"]
  HTML --> X{"Extractor"}
  X -->|"no model"| API["APIExtractor -> api.parsera.org"]
  X -->|"model"| CH["ChunksTabularExtractor"]
  X -->|"model, typed=True"| ST["StructuredExtractor"]
  CH --> MD["markdownify -> chunks"]
  ST --> MD
  MD --> LLM["LangChain chat model"]
  LLM --> J["JSON records"]
  API --> J
Component Path Role
Facade parsera/parsera.py Parsera picks an extractor, owns a PageLoader, exposes run (sync) and arun (async)
Page loader parsera/page.py Firefox launch, context with proxy and cookies, playwright-stealth, scrolling, iframe capture
API extractor parsera/engine/api_extractor.py Extractor base class and the hosted-API client
Local extractors parsera/engine/simple_extractor.py LocalExtractor plus TabularExtractor, ListExtractor, ItemExtractor (prompt variants)
Chunking parsera/engine/chunks_extractor.py Token-based splitting, sequential append prompts, final merge call
Typed output parsera/engine/structured_extractor.py Builds a Pydantic schema from field types and uses with_structured_output
Model wrappers parsera/engine/model.py GPT4oMiniModel, AzureModel, HuggingFaceModel singletons
CLI parsera/main.py python -m parsera.main URL --scheme/--file --scrolls --output

How a request flows

Take Parsera(model=llm).run(url, elements):

  1. Choose the extractor. The constructor accepts either a model or an extractor, never both. No model means APIExtractor. A model means ChunksTabularExtractor, or StructuredExtractor when typed=True (parsera.py).
  2. Open a session once. _run creates the browser session on the first call only: PageLoader.create_session launches Firefox headless, opens a context with the optional proxy and cookies, and, with stealth=True (the default), rebuilds the context and applies stealth_async. An initial_script callback runs here, which is where login flows go (parsera.py, page.py).
  3. Fetch. fetch_page calls page.goto, waits for networkidle (a timeout is silently ignored), runs the per-call playwright_script, then either scrolls or reads the full HTML (page.py).
  4. Collect HTML. get_full_html concatenates the main document’s outerHTML with every non-detached iframe’s HTML, fetched concurrently with asyncio.gather (page.py).
  5. Convert and split. The extractor runs MarkdownConverter().convert(html), then splits the Markdown with LangChain’s RecursiveCharacterTextSplitter, measuring length in o200k_base tokens. The default chunk is 100,000 tokens with a one-third overlap (chunks_extractor.py).
  6. Extract. One chunk means one call: a few-shot system prompt asking for a list of flat records, plus the Markdown and the requested fields, parsed with JsonOutputParser (chunks_extractor.py).
  7. Multi-chunk path. For several chunks, the loop runs sequentially. Each call gets the tail of the previous chunk’s records and is asked to fix truncated rows and continue the sequence. A final LLM call merges all per-chunk lists and removes duplicates (chunks_extractor.py). A page split into N chunks costs N + 1 calls.
  8. Return. The parsed JSON (normally list[dict]) is returned as is. There is no schema validation on this path.

Key components

PageLoader

PageLoader holds one Playwright instance, one browser, one context and one page, and reuses them across calls on the same Parsera object. The browser is always Firefox. Stealth mode reads navigator.userAgent, replaces HeadlessChrome/ with Chrome/, reopens the context with that UA, and runs playwright-stealth with its UA, plugins and vendor patches turned off (page.py). The UA rewrite targets Chrome’s headless marker, so it does nothing on Firefox. The remaining playwright-stealth scripts are the actual evasion layer. There is no CAPTCHA handling and no rate limiting.

Infinite scroll

scroll_page installs a MutationObserver that records the outerHTML of every removed element. It then scrolls to the bottom up to scrolls_limit times, sleeping 2 s each time and stopping when the height stops growing. It returns every captured snapshot plus the removed nodes, concatenated (page.py). This handles virtualised lists that drop off-screen rows. The cost is heavy duplication, because each snapshot repeats the whole page, and the merge step is left to the LLM. Iframes are not captured on this path.

Extractors

All local extractors share LocalExtractor.run: convert to Markdown, fill a template, call model.ainvoke, parse JSON (simple_extractor.py). The subclasses differ only in system prompt: TabularExtractor (list of rows), ListExtractor (dict of lists) and ItemExtractor (one dict). These single-call extractors never chunk. ChunksTabularExtractor is the only one Parsera builds for you.

Typed extraction

With typed=True, each field is {"description": ..., "type": "string"|"integer"|"number"|"bool"|"list"|"object"|"any"}. StructuredExtractor.create_schema builds a Pydantic record model where every field is optional, wraps it in {reasoning, data: [record]}, and binds it with model.with_structured_output (structured_extractor.py). The reasoning field is a built-in chain-of-thought slot that costs output tokens and is then dropped. An all-null result is turned into [].

Hosted API path

APIExtractor.run posts the raw HTML (not Markdown), the prompt, the attributes as {name, description} pairs and mode: "standard" to the Parsera API with a 180 s timeout. A detail key in the response becomes a Python warning, not an exception (api_extractor.py). It uses blocking requests inside an async method, so arun calls on this path block the event loop.

Extending it

  • Any LangChain model. Pass model= with an OpenAI, Azure, Ollama or other BaseChatModel. HuggingFaceModel wraps a local transformers pipeline and runs it in a thread executor (model.py). Note that the bundled wrappers are @singletons, so the first constructor arguments win for the whole process.
  • Custom extractor. Subclass Extractor (one async run(content, attributes, prompt)) or a LocalExtractor with your own system_prompt, and pass it as extractor=. ChunksTabularExtractor also accepts chunk_size, a token_counter and a custom MarkdownConverter.
  • Playwright hooks. initial_script runs once per session, and playwright_script runs after every navigation. Both receive and return a Page, so clicks, logins and waits go there.
  • Direct loader use. PageLoader(browser=...) accepts your own Playwright browser, for example a headed one with slow_mo for debugging, as the bundled scrolling example does.

Running it

  • pip install parsera and playwright install (Firefox is the browser that is used). Python 3.10 or newer.
  • Hosted mode: set PARSERA_API_KEY and call Parsera().run(url=..., elements=...).
  • Local mode: pass model=, and set whatever key that provider needs.
  • CLI: python -m parsera.main URL --scheme '{"title":"h1"}' --scrolls 5 --output out.json. The CLI always builds GPT4oMiniModel, so it needs OPENAI_API_KEY, not a Parsera key (main.py).
  • Docker: the image installs Poetry and Playwright Firefox and uses the CLI as its entrypoint. docker-compose.yaml mounts a scheme.json and an output/ folder.

Strengths and caveats

  • Strength: tiny and readable. The whole data path fits in five files. It is easy to audit, fork, or lift the chunk-and-merge idea from.
  • Strength: chunking that respects rows. Overlapping chunks, “continue this sequence” prompts and a merge call are a reasonable answer to tables longer than a context window.
  • Strength: lazy-load capture. The removed-node observer recovers rows that virtualised lists delete while scrolling, which many scrapers silently miss.
  • Caveat: the default is a paid cloud call. Parsera() with no model sends the whole page HTML to a third-party API. You must pass model= to keep data local.
  • Caveat: full LLM cost on every page. Nothing is cached or turned into selectors. A large page that splits into N chunks costs N + 1 sequential calls, and the merge call re-sends all extracted data.
  • Caveat: weak anti-bot. Firefox plus a Chrome-oriented UA patch plus playwright-stealth is basic. There are no CAPTCHA, retry or backoff features.
  • Caveat: session reuse is fragile. run() calls asyncio.run each time, while the Playwright session from the first call is kept on the loader. Repeated sync calls on one instance can hit a session bound to a closed event loop. Use arun for multi-page jobs, or a fresh instance per URL.
  • Caveat: hygiene. The committed run.py sets a literal PARSERA_API_KEY value, and the test suite needs live Azure OpenAI credentials and network access.

Sources: code at 865d777, deepwiki-open wiki (10 pages), OpenDeepWiki wiki (9 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 865d777. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Pages are fetched exclusively via Playwright controlling a headless Firefox browser — there is no plain-HTTP fallback. The PageLoader class (parsera/page.py:19) manages the Playwright lifecycle: it launches Firefox headless (page.py:37), creates a browser context, opens a page, navigates to the URL with page.goto(), and waits for a load state (default "networkidle", also accepts "domcontentloaded" or "load", with a try/except pass on timeout at page.py:183-190). JavaScript rendering is therefore fully supported — all JS on the page executes before content is captured. For JS-heavy or infinite-scroll pages, the scroll_page() method (page.py:92-144) scrolls programmatically with configurable scroll-limit, a 2-second sleep between scrolls, and a MutationObserver that captures elements removed during lazy-loading. Content is retrieved as raw HTML via document.documentElement.outerHTML, including content from all iframes fetched in parallel (page.py:155-173). There is no support for PDF or image content — the library reads HTML only. A custom Playwright script can be injected at session creation (initial_script) or before each page fetch (playwright_script), enabling auth flows, cookie injection, or custom wait logic (page.py:78-90, parsera.py:60-71).

How is content extracted or converted?

answered

Extraction follows a two-stage pipeline: HTML→Markdown conversion via markdownify (MarkdownConverter), then LLM-based structured parsing. Every extractor in parsera/engine/simple_extractor.py inherits from LocalExtractor and calls self.converter.convert(content) before sending to the LLM (simple_extractor.py:68). There are no readability-style heuristics or boilerplate-removal rules — the full page markdown is sent to the model, which is expected to ignore irrelevant content. Five extractor types exist: TabularExtractor returns a list-of-dicts ([{...}, {...}]) for table-like data (simple_extractor.py:156-157). ListExtractor returns dict-of-lists (simple_extractor.py:202-203). ItemExtractor returns a single dict (simple_extractor.py:248-249). ChunksTabularExtractor (chunks_extractor.py:184) extends TabularExtractor with recursive character text-splitting (LangChain's RecursiveCharacterTextSplitter with ~33% overlap, chunk_size default 100k tokens) for pages that exceed context windows; it extracts each chunk sequentially, passing previous results as context to handle truncated rows, then merges all chunk results via a second LLM call (chunks_extractor.py:274-292). StructuredExtractor (structured_extractor.py:31) extends ChunksTabularExtractor with Pydantic schema-based structured output: it dynamically creates a pydantic.BaseModel from user-supplied type/description fields and calls model.with_structured_output() (structured_extractor.py:102-142). The user specifies fields as a dict of field_name → description (Tabular) or field_name → {type, description} (Structured).

How are LLMs used, if at all?

answered

LLMs are the core extraction engine — every extractor sends the full page markdown (or chunks thereof) to a LangChain BaseChatModel (langchain_core.language_models) via await model.ainvoke(messages), using a SystemMessage (the extractor's behaviour prompt with few-shot examples) + HumanMessage (the markdown content and field definitions). The default model when none is provided is GPT-4o-mini (model.py:36-44), wrapped as a singleton GPT4oMiniModel. Alternatively users can pass any LangChain-compatible chat model: Azure OpenAI via AzureModel (model.py:20-32), Ollama models (documented in docs/features/custom-models.md), or HuggingFace pipelines via HuggingFaceModel (model.py:48-112), which runs the pipeline in a thread executor since it lacks native async. Chunking is used for large pages: ChunksTabularExtractor uses RecursiveCharacterTextSplitter with a default chunk_size=100000 tokens (chunks_extractor.py:227-231), counted via tiktoken with the o200k_base encoding. When chunked, each chunk first extends/continues the previous extraction via an "append" prompt, then a final LLM call merges all chunk results deduplicating by proximity to boundaries. For structured output, the StructuredExtractor calls model.with_structured_output(schema=OutputSchema) on a dynamically-generated Pydantic model (structured_extractor.py:142). There is no explicit cost control mechanism beyond the user choosing a smaller model or limiting chunk size. The APIExtractor (api_extractor.py:18-51) bypasses local LLM usage entirely, sending the raw HTML content to https://api.parsera.org/v1/parse with a PARSERA_API_KEY, making the cloud service the LLM backend.

Editor's note. Correction: with no model and no extractor, Parsera() uses APIExtractor, which sends the raw page HTML to the hosted api.parsera.org endpoint with PARSERA_API_KEY (parsera/parsera.py L36-L37). GPT4oMiniModel is only the default of the CLI in parsera/main.py; library users must pass model= to run extraction locally.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-bot evasion is handled through a dedicated stealth mode (enabled by default) using the playwright-stealth library. In PageLoader.stealth() (page.py:40-66), after the initial context is created and a page opened, the library replaces the User-Agent string to remove the HeadlessChrome/ marker — it evaluates navigator.userAgent, strips HeadlessChrome/ down to Chrome/ (page.py:44-45), then closes the original context and creates a new one with the spoofed User-Agent. It then applies stealth_async() with a StealthConfig that skips overriding navigator_user_agent, navigator_plugins, and navigator_vendor (since the UA is already patched manually, page.py:57-63). Proxy support is built-in: a ProxySettings TypedDict (server, optional bypass/username/password) can be passed to run()/arun(), and when creating the Playwright context, proxy=proxy_settings is forwarded to Playwright's new_context() (page.py:48-49, page.py:76). CAPTCHA handling is not implemented in the codebase. Rate limiting is not implemented — there is no throttle, backoff, or concurrency limiter on page fetches. Cookie injection is supported: custom cookies passed to the constructor (custom_cookies) are added to the browser context via context.add_cookies() (page.py:51-55, page.py:78-82) with validation that raises CookiesValidationException on failure. The browser is always headless Firefox (page.py:37), though examples show headless=False is possible with a custom PageLoader or by passing a custom browser.

How is crawling at scale implemented?

not applicable

Parsera does not implement crawling at scale. It is a single-page scraping library with no queue system, no URL deduplication, no depth or crawl-limit parameters, no robots.txt parsing, no politeness delays, and no distributed worker support. Each run()/arun() call fetches exactly one URL, extracts its data, and returns — the user must build any multi-page or crawl logic externally. There is no concurrency model for page fetching (Playwright handles one page at a time per session). The only "scale" feature is page-level chunking for long content, which addresses LLM context limits, not crawling breadth.

What is the developer interface?

answered

The primary interface is the Python library API via the Parsera class (parsera/parsera.py:13-115). Instantiate with optional model (LangChain chat model), extractor, initial_script (Playwright callback), stealth (bool), and custom_cookies. Call .run(url, elements, prompt, ...) for synchronous usage (wraps asyncio.run()) or .arun(...) for async. The result is a list[dict] (for tabular extractors) or dict (for Item/List extractors). A CLI is provided via python -m parsera.main (main.py:56-181), accepting a URL, --scheme (JSON string of field definitions), --file (JSON file with same), --scrolls (int), and --output (file path, default "output.json"). The CLI uses GPT4oMiniModel by default and requires OPENAI_API_KEY or PARSERA_API_KEY in the environment. A Docker entrypoint is configured (Dockerfile:42) that runs the CLI with environment variables for URL, FILE, and OUTPUT. There is no REST API service served by the project (the APIExtractor is a client to a hosted API at api.parsera.org, not a server). There is no MCP server, no web UI, and no graphical interface. Output formats are limited to Python dicts (in-code) and JSON files (CLI). Language bindings are Python-only. The project exposes no language-agnostic protocol.

Editor's note. Correction: the CLI always constructs GPT4oMiniModel (parsera/main.py L121-L135), so it needs OPENAI_API_KEY; a PARSERA_API_KEY alone does not make the CLI work.