raznem/parsera
Python library that renders a page in headless Firefox, converts it to Markdown and has an LLM return the fields you describe as JSON.
Overview
Parsera is a small Python library for “describe the fields, get JSON back” scraping. You give it a URL and a dict such as {"Title": "News title", "Points": "Number of points"}. It opens the page in headless Firefox through Playwright, grabs the rendered HTML (iframes included), turns it into Markdown with markdownify, and asks a language model to return a list of records with those keys. There are no selectors, no XPath and no learned extraction rules. Every run re-reads the page with the model.
The whole package is about 1,100 lines of code. The interesting decision is where the LLM lives. A bare Parsera() does not call OpenAI at all. It sends the page to the vendor’s hosted endpoint, https://api.parsera.org/v1/parse, authenticated with PARSERA_API_KEY (parsera.py, api_extractor.py). Pass any LangChain chat model and extraction runs locally instead, through a chunking extractor with its own prompts. So the open-source library is both a client for the commercial Parsera service and a self-contained LLM scraper.
It suits one-off or low-volume jobs where the page layout changes often and per-page LLM cost is acceptable: a listing page, a product table, a news front page. It is a single-URL tool. There is no crawler, queue, link following or retry logic.
Architecture
flowchart LR
U["Your code / CLI"] --> P["Parsera.run / arun"]
P --> L["PageLoader"]
L --> FF["Headless Firefox (Playwright)"]
L --> HTML["HTML + iframe HTML"]
HTML --> X{"Extractor"}
X -->|"no model"| API["APIExtractor -> api.parsera.org"]
X -->|"model"| CH["ChunksTabularExtractor"]
X -->|"model, typed=True"| ST["StructuredExtractor"]
CH --> MD["markdownify -> chunks"]
ST --> MD
MD --> LLM["LangChain chat model"]
LLM --> J["JSON records"]
API --> J
| Component | Path | Role |
|---|---|---|
| Facade | parsera/parsera.py |
Parsera picks an extractor, owns a PageLoader, exposes run (sync) and arun (async) |
| Page loader | parsera/page.py |
Firefox launch, context with proxy and cookies, playwright-stealth, scrolling, iframe capture |
| API extractor | parsera/engine/api_extractor.py |
Extractor base class and the hosted-API client |
| Local extractors | parsera/engine/simple_extractor.py |
LocalExtractor plus TabularExtractor, ListExtractor, ItemExtractor (prompt variants) |
| Chunking | parsera/engine/chunks_extractor.py |
Token-based splitting, sequential append prompts, final merge call |
| Typed output | parsera/engine/structured_extractor.py |
Builds a Pydantic schema from field types and uses with_structured_output |
| Model wrappers | parsera/engine/model.py |
GPT4oMiniModel, AzureModel, HuggingFaceModel singletons |
| CLI | parsera/main.py |
python -m parsera.main URL --scheme/--file --scrolls --output |
How a request flows
Take Parsera(model=llm).run(url, elements):
- Choose the extractor. The constructor accepts either a
modelor anextractor, never both. No model meansAPIExtractor. A model meansChunksTabularExtractor, orStructuredExtractorwhentyped=True(parsera.py). - Open a session once.
_runcreates the browser session on the first call only:PageLoader.create_sessionlaunches Firefox headless, opens a context with the optional proxy and cookies, and, withstealth=True(the default), rebuilds the context and appliesstealth_async. Aninitial_scriptcallback runs here, which is where login flows go (parsera.py, page.py). - Fetch.
fetch_pagecallspage.goto, waits fornetworkidle(a timeout is silently ignored), runs the per-callplaywright_script, then either scrolls or reads the full HTML (page.py). - Collect HTML.
get_full_htmlconcatenates the main document’souterHTMLwith every non-detached iframe’s HTML, fetched concurrently withasyncio.gather(page.py). - Convert and split. The extractor runs
MarkdownConverter().convert(html), then splits the Markdown with LangChain’sRecursiveCharacterTextSplitter, measuring length ino200k_basetokens. The default chunk is 100,000 tokens with a one-third overlap (chunks_extractor.py). - Extract. One chunk means one call: a few-shot system prompt asking for a list of flat records, plus the Markdown and the requested fields, parsed with
JsonOutputParser(chunks_extractor.py). - Multi-chunk path. For several chunks, the loop runs sequentially. Each call gets the tail of the previous chunk’s records and is asked to fix truncated rows and continue the sequence. A final LLM call merges all per-chunk lists and removes duplicates (chunks_extractor.py). A page split into N chunks costs N + 1 calls.
- Return. The parsed JSON (normally
list[dict]) is returned as is. There is no schema validation on this path.
Key components
PageLoader
PageLoader holds one Playwright instance, one browser, one context and one page, and reuses them across calls on the same Parsera object. The browser is always Firefox. Stealth mode reads navigator.userAgent, replaces HeadlessChrome/ with Chrome/, reopens the context with that UA, and runs playwright-stealth with its UA, plugins and vendor patches turned off (page.py). The UA rewrite targets Chrome’s headless marker, so it does nothing on Firefox. The remaining playwright-stealth scripts are the actual evasion layer. There is no CAPTCHA handling and no rate limiting.
Infinite scroll
scroll_page installs a MutationObserver that records the outerHTML of every removed element. It then scrolls to the bottom up to scrolls_limit times, sleeping 2 s each time and stopping when the height stops growing. It returns every captured snapshot plus the removed nodes, concatenated (page.py). This handles virtualised lists that drop off-screen rows. The cost is heavy duplication, because each snapshot repeats the whole page, and the merge step is left to the LLM. Iframes are not captured on this path.
Extractors
All local extractors share LocalExtractor.run: convert to Markdown, fill a template, call model.ainvoke, parse JSON (simple_extractor.py). The subclasses differ only in system prompt: TabularExtractor (list of rows), ListExtractor (dict of lists) and ItemExtractor (one dict). These single-call extractors never chunk. ChunksTabularExtractor is the only one Parsera builds for you.
Typed extraction
With typed=True, each field is {"description": ..., "type": "string"|"integer"|"number"|"bool"|"list"|"object"|"any"}. StructuredExtractor.create_schema builds a Pydantic record model where every field is optional, wraps it in {reasoning, data: [record]}, and binds it with model.with_structured_output (structured_extractor.py). The reasoning field is a built-in chain-of-thought slot that costs output tokens and is then dropped. An all-null result is turned into [].
Hosted API path
APIExtractor.run posts the raw HTML (not Markdown), the prompt, the attributes as {name, description} pairs and mode: "standard" to the Parsera API with a 180 s timeout. A detail key in the response becomes a Python warning, not an exception (api_extractor.py). It uses blocking requests inside an async method, so arun calls on this path block the event loop.
Extending it
- Any LangChain model. Pass
model=with an OpenAI, Azure, Ollama or otherBaseChatModel.HuggingFaceModelwraps a localtransformerspipeline and runs it in a thread executor (model.py). Note that the bundled wrappers are@singletons, so the first constructor arguments win for the whole process. - Custom extractor. Subclass
Extractor(one asyncrun(content, attributes, prompt)) or aLocalExtractorwith your ownsystem_prompt, and pass it asextractor=.ChunksTabularExtractoralso acceptschunk_size, atoken_counterand a customMarkdownConverter. - Playwright hooks.
initial_scriptruns once per session, andplaywright_scriptruns after every navigation. Both receive and return aPage, so clicks, logins and waits go there. - Direct loader use.
PageLoader(browser=...)accepts your own Playwright browser, for example a headed one withslow_mofor debugging, as the bundled scrolling example does.
Running it
pip install parseraandplaywright install(Firefox is the browser that is used). Python 3.10 or newer.- Hosted mode: set
PARSERA_API_KEYand callParsera().run(url=..., elements=...). - Local mode: pass
model=, and set whatever key that provider needs. - CLI:
python -m parsera.main URL --scheme '{"title":"h1"}' --scrolls 5 --output out.json. The CLI always buildsGPT4oMiniModel, so it needsOPENAI_API_KEY, not a Parsera key (main.py). - Docker: the image installs Poetry and Playwright Firefox and uses the CLI as its entrypoint.
docker-compose.yamlmounts ascheme.jsonand anoutput/folder.
Strengths and caveats
- Strength: tiny and readable. The whole data path fits in five files. It is easy to audit, fork, or lift the chunk-and-merge idea from.
- Strength: chunking that respects rows. Overlapping chunks, “continue this sequence” prompts and a merge call are a reasonable answer to tables longer than a context window.
- Strength: lazy-load capture. The removed-node observer recovers rows that virtualised lists delete while scrolling, which many scrapers silently miss.
- Caveat: the default is a paid cloud call.
Parsera()with no model sends the whole page HTML to a third-party API. You must passmodel=to keep data local. - Caveat: full LLM cost on every page. Nothing is cached or turned into selectors. A large page that splits into N chunks costs N + 1 sequential calls, and the merge call re-sends all extracted data.
- Caveat: weak anti-bot. Firefox plus a Chrome-oriented UA patch plus
playwright-stealthis basic. There are no CAPTCHA, retry or backoff features. - Caveat: session reuse is fragile.
run()callsasyncio.runeach time, while the Playwright session from the first call is kept on the loader. Repeated sync calls on one instance can hit a session bound to a closed event loop. Usearunfor multi-page jobs, or a fresh instance per URL. - Caveat: hygiene. The committed
run.pysets a literalPARSERA_API_KEYvalue, and the test suite needs live Azure OpenAI credentials and network access.
Sources: code at 865d777, deepwiki-open wiki (10 pages), OpenDeepWiki wiki (9 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 865d777. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredPages are fetched exclusively via Playwright controlling a headless Firefox browser — there is no plain-HTTP fallback. The PageLoader class (parsera/page.py:19) manages the Playwright lifecycle: it launches Firefox headless (page.py:37), creates a browser context, opens a page, navigates to the URL with page.goto(), and waits for a load state (default "networkidle", also accepts "domcontentloaded" or "load", with a try/except pass on timeout at page.py:183-190). JavaScript rendering is therefore fully supported — all JS on the page executes before content is captured. For JS-heavy or infinite-scroll pages, the scroll_page() method (page.py:92-144) scrolls programmatically with configurable scroll-limit, a 2-second sleep between scrolls, and a MutationObserver that captures elements removed during lazy-loading. Content is retrieved as raw HTML via document.documentElement.outerHTML, including content from all iframes fetched in parallel (page.py:155-173). There is no support for PDF or image content — the library reads HTML only. A custom Playwright script can be injected at session creation (initial_script) or before each page fetch (playwright_script), enabling auth flows, cookie injection, or custom wait logic (page.py:78-90, parsera.py:60-71).
How is content extracted or converted?
answeredExtraction follows a two-stage pipeline: HTML→Markdown conversion via markdownify (MarkdownConverter), then LLM-based structured parsing. Every extractor in parsera/engine/simple_extractor.py inherits from LocalExtractor and calls self.converter.convert(content) before sending to the LLM (simple_extractor.py:68). There are no readability-style heuristics or boilerplate-removal rules — the full page markdown is sent to the model, which is expected to ignore irrelevant content. Five extractor types exist: TabularExtractor returns a list-of-dicts ([{...}, {...}]) for table-like data (simple_extractor.py:156-157). ListExtractor returns dict-of-lists (simple_extractor.py:202-203). ItemExtractor returns a single dict (simple_extractor.py:248-249). ChunksTabularExtractor (chunks_extractor.py:184) extends TabularExtractor with recursive character text-splitting (LangChain's RecursiveCharacterTextSplitter with ~33% overlap, chunk_size default 100k tokens) for pages that exceed context windows; it extracts each chunk sequentially, passing previous results as context to handle truncated rows, then merges all chunk results via a second LLM call (chunks_extractor.py:274-292). StructuredExtractor (structured_extractor.py:31) extends ChunksTabularExtractor with Pydantic schema-based structured output: it dynamically creates a pydantic.BaseModel from user-supplied type/description fields and calls model.with_structured_output() (structured_extractor.py:102-142). The user specifies fields as a dict of field_name → description (Tabular) or field_name → {type, description} (Structured).
How are LLMs used, if at all?
answeredLLMs are the core extraction engine — every extractor sends the full page markdown (or chunks thereof) to a LangChain BaseChatModel (langchain_core.language_models) via await model.ainvoke(messages), using a SystemMessage (the extractor's behaviour prompt with few-shot examples) + HumanMessage (the markdown content and field definitions). The default model when none is provided is GPT-4o-mini (model.py:36-44), wrapped as a singleton GPT4oMiniModel. Alternatively users can pass any LangChain-compatible chat model: Azure OpenAI via AzureModel (model.py:20-32), Ollama models (documented in docs/features/custom-models.md), or HuggingFace pipelines via HuggingFaceModel (model.py:48-112), which runs the pipeline in a thread executor since it lacks native async. Chunking is used for large pages: ChunksTabularExtractor uses RecursiveCharacterTextSplitter with a default chunk_size=100000 tokens (chunks_extractor.py:227-231), counted via tiktoken with the o200k_base encoding. When chunked, each chunk first extends/continues the previous extraction via an "append" prompt, then a final LLM call merges all chunk results deduplicating by proximity to boundaries. For structured output, the StructuredExtractor calls model.with_structured_output(schema=OutputSchema) on a dynamically-generated Pydantic model (structured_extractor.py:142). There is no explicit cost control mechanism beyond the user choosing a smaller model or limiting chunk size. The APIExtractor (api_extractor.py:18-51) bypasses local LLM usage entirely, sending the raw HTML content to https://api.parsera.org/v1/parse with a PARSERA_API_KEY, making the cloud service the LLM backend.
model and no extractor, Parsera() uses APIExtractor, which sends the raw page HTML to the hosted api.parsera.org endpoint with PARSERA_API_KEY (parsera/parsera.py L36-L37). GPT4oMiniModel is only the default of the CLI in parsera/main.py; library users must pass model= to run extraction locally.How are anti-bot measures, proxies and fingerprinting handled?
answeredAnti-bot evasion is handled through a dedicated stealth mode (enabled by default) using the playwright-stealth library. In PageLoader.stealth() (page.py:40-66), after the initial context is created and a page opened, the library replaces the User-Agent string to remove the HeadlessChrome/ marker — it evaluates navigator.userAgent, strips HeadlessChrome/ down to Chrome/ (page.py:44-45), then closes the original context and creates a new one with the spoofed User-Agent. It then applies stealth_async() with a StealthConfig that skips overriding navigator_user_agent, navigator_plugins, and navigator_vendor (since the UA is already patched manually, page.py:57-63). Proxy support is built-in: a ProxySettings TypedDict (server, optional bypass/username/password) can be passed to run()/arun(), and when creating the Playwright context, proxy=proxy_settings is forwarded to Playwright's new_context() (page.py:48-49, page.py:76). CAPTCHA handling is not implemented in the codebase. Rate limiting is not implemented — there is no throttle, backoff, or concurrency limiter on page fetches. Cookie injection is supported: custom cookies passed to the constructor (custom_cookies) are added to the browser context via context.add_cookies() (page.py:51-55, page.py:78-82) with validation that raises CookiesValidationException on failure. The browser is always headless Firefox (page.py:37), though examples show headless=False is possible with a custom PageLoader or by passing a custom browser.
How is crawling at scale implemented?
not applicableParsera does not implement crawling at scale. It is a single-page scraping library with no queue system, no URL deduplication, no depth or crawl-limit parameters, no robots.txt parsing, no politeness delays, and no distributed worker support. Each run()/arun() call fetches exactly one URL, extracts its data, and returns — the user must build any multi-page or crawl logic externally. There is no concurrency model for page fetching (Playwright handles one page at a time per session). The only "scale" feature is page-level chunking for long content, which addresses LLM context limits, not crawling breadth.
What is the developer interface?
answeredThe primary interface is the Python library API via the Parsera class (parsera/parsera.py:13-115). Instantiate with optional model (LangChain chat model), extractor, initial_script (Playwright callback), stealth (bool), and custom_cookies. Call .run(url, elements, prompt, ...) for synchronous usage (wraps asyncio.run()) or .arun(...) for async. The result is a list[dict] (for tabular extractors) or dict (for Item/List extractors). A CLI is provided via python -m parsera.main (main.py:56-181), accepting a URL, --scheme (JSON string of field definitions), --file (JSON file with same), --scrolls (int), and --output (file path, default "output.json"). The CLI uses GPT4oMiniModel by default and requires OPENAI_API_KEY or PARSERA_API_KEY in the environment. A Docker entrypoint is configured (Dockerfile:42) that runs the CLI with environment variables for URL, FILE, and OUTPUT. There is no REST API service served by the project (the APIExtractor is a client to a hosted API at api.parsera.org, not a server). There is no MCP server, no web UI, and no graphical interface. Output formats are limited to Python dicts (in-code) and JSON files (CLI). Language bindings are Python-only. The project exposes no language-agnostic protocol.
GPT4oMiniModel (parsera/main.py L121-L135), so it needs OPENAI_API_KEY; a PARSERA_API_KEY alone does not make the CLI work.