# raznem/parsera

> Python library that renders a page in headless Firefox, converts it to Markdown and has an LLM return the fields you describe as JSON.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/raznem/parsera (reviewed at commit `865d7772fc92af9dfe3622ca1adcd5e5b7afd07f`, 2025-12-17)
- Stars: 1366 · Language: Python · License: GPL-2.0
- Canonical page: https://llms-technical-reviews.com/p/parsera/

## Overview

Parsera is a small Python library for "describe the fields, get JSON back" scraping. You give it a URL and a dict such as `{"Title": "News title", "Points": "Number of points"}`. It opens the page in headless Firefox through Playwright, grabs the rendered HTML (iframes included), turns it into Markdown with `markdownify`, and asks a language model to return a list of records with those keys. There are no selectors, no XPath and no learned extraction rules. Every run re-reads the page with the model.

The whole package is about 1,100 lines of code. The interesting decision is where the LLM lives. A bare `Parsera()` does not call OpenAI at all. It sends the page to the vendor's hosted endpoint, `https://api.parsera.org/v1/parse`, authenticated with `PARSERA_API_KEY` ([parsera.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/parsera.py#L36-L49), [api_extractor.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/api_extractor.py#L7-L51)). Pass any LangChain chat model and extraction runs locally instead, through a chunking extractor with its own prompts. So the open-source library is both a client for the commercial Parsera service and a self-contained LLM scraper.

It suits one-off or low-volume jobs where the page layout changes often and per-page LLM cost is acceptable: a listing page, a product table, a news front page. It is a single-URL tool. There is no crawler, queue, link following or retry logic.

## Architecture

```mermaid
flowchart LR
  U["Your code / CLI"] --> P["Parsera.run / arun"]
  P --> L["PageLoader"]
  L --> FF["Headless Firefox (Playwright)"]
  L --> HTML["HTML + iframe HTML"]
  HTML --> X{"Extractor"}
  X -->|"no model"| API["APIExtractor -> api.parsera.org"]
  X -->|"model"| CH["ChunksTabularExtractor"]
  X -->|"model, typed=True"| ST["StructuredExtractor"]
  CH --> MD["markdownify -> chunks"]
  ST --> MD
  MD --> LLM["LangChain chat model"]
  LLM --> J["JSON records"]
  API --> J
```

| Component | Path | Role |
|---|---|---|
| Facade | `parsera/parsera.py` | `Parsera` picks an extractor, owns a `PageLoader`, exposes `run` (sync) and `arun` (async) |
| Page loader | `parsera/page.py` | Firefox launch, context with proxy and cookies, `playwright-stealth`, scrolling, iframe capture |
| API extractor | `parsera/engine/api_extractor.py` | `Extractor` base class and the hosted-API client |
| Local extractors | `parsera/engine/simple_extractor.py` | `LocalExtractor` plus `TabularExtractor`, `ListExtractor`, `ItemExtractor` (prompt variants) |
| Chunking | `parsera/engine/chunks_extractor.py` | Token-based splitting, sequential append prompts, final merge call |
| Typed output | `parsera/engine/structured_extractor.py` | Builds a Pydantic schema from field types and uses `with_structured_output` |
| Model wrappers | `parsera/engine/model.py` | `GPT4oMiniModel`, `AzureModel`, `HuggingFaceModel` singletons |
| CLI | `parsera/main.py` | `python -m parsera.main URL --scheme/--file --scrolls --output` |

## How a request flows

Take `Parsera(model=llm).run(url, elements)`:

1. **Choose the extractor.** The constructor accepts either a `model` or an `extractor`, never both. No model means `APIExtractor`. A model means `ChunksTabularExtractor`, or `StructuredExtractor` when `typed=True` ([parsera.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/parsera.py#L14-L52)).
2. **Open a session once.** `_run` creates the browser session on the first call only: `PageLoader.create_session` launches Firefox headless, opens a context with the optional proxy and cookies, and, with `stealth=True` (the default), rebuilds the context and applies `stealth_async`. An `initial_script` callback runs here, which is where login flows go ([parsera.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/parsera.py#L54-L77), [page.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L29-L90)).
3. **Fetch.** `fetch_page` calls `page.goto`, waits for `networkidle` (a timeout is silently ignored), runs the per-call `playwright_script`, then either scrolls or reads the full HTML ([page.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L175-L201)).
4. **Collect HTML.** `get_full_html` concatenates the main document's `outerHTML` with every non-detached iframe's HTML, fetched concurrently with `asyncio.gather` ([page.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L146-L173)).
5. **Convert and split.** The extractor runs `MarkdownConverter().convert(html)`, then splits the Markdown with LangChain's `RecursiveCharacterTextSplitter`, measuring length in `o200k_base` tokens. The default chunk is 100,000 tokens with a one-third overlap ([chunks_extractor.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/chunks_extractor.py#L191-L232)).
6. **Extract.** One chunk means one call: a few-shot system prompt asking for a list of flat records, plus the Markdown and the requested fields, parsed with `JsonOutputParser` ([chunks_extractor.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/chunks_extractor.py#L239-L272)).
7. **Multi-chunk path.** For several chunks, the loop runs sequentially. Each call gets the tail of the previous chunk's records and is asked to fix truncated rows and continue the sequence. A final LLM call merges all per-chunk lists and removes duplicates ([chunks_extractor.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/chunks_extractor.py#L274-L331)). A page split into N chunks costs N + 1 calls.
8. **Return.** The parsed JSON (normally `list[dict]`) is returned as is. There is no schema validation on this path.

## Key components

### PageLoader

`PageLoader` holds one Playwright instance, one browser, one context and one page, and reuses them across calls on the same `Parsera` object. The browser is always Firefox. Stealth mode reads `navigator.userAgent`, replaces `HeadlessChrome/` with `Chrome/`, reopens the context with that UA, and runs `playwright-stealth` with its UA, plugins and vendor patches turned off ([page.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L39-L66)). The UA rewrite targets Chrome's headless marker, so it does nothing on Firefox. The remaining `playwright-stealth` scripts are the actual evasion layer. There is no CAPTCHA handling and no rate limiting.

### Infinite scroll

`scroll_page` installs a `MutationObserver` that records the `outerHTML` of every removed element. It then scrolls to the bottom up to `scrolls_limit` times, sleeping 2 s each time and stopping when the height stops growing. It returns every captured snapshot plus the removed nodes, concatenated ([page.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L92-L144)). This handles virtualised lists that drop off-screen rows. The cost is heavy duplication, because each snapshot repeats the whole page, and the merge step is left to the LLM. Iframes are not captured on this path.

### Extractors

All local extractors share `LocalExtractor.run`: convert to Markdown, fill a template, call `model.ainvoke`, parse JSON ([simple_extractor.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/simple_extractor.py#L39-L86)). The subclasses differ only in system prompt: `TabularExtractor` (list of rows), `ListExtractor` (dict of lists) and `ItemExtractor` (one dict). These single-call extractors never chunk. `ChunksTabularExtractor` is the only one `Parsera` builds for you.

### Typed extraction

With `typed=True`, each field is `{"description": ..., "type": "string"|"integer"|"number"|"bool"|"list"|"object"|"any"}`. `StructuredExtractor.create_schema` builds a Pydantic record model where every field is optional, wraps it in `{reasoning, data: [record]}`, and binds it with `model.with_structured_output` ([structured_extractor.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/structured_extractor.py#L102-L146)). The `reasoning` field is a built-in chain-of-thought slot that costs output tokens and is then dropped. An all-null result is turned into `[]`.

### Hosted API path

`APIExtractor.run` posts the raw HTML (not Markdown), the prompt, the attributes as `{name, description}` pairs and `mode: "standard"` to the Parsera API with a 180 s timeout. A `detail` key in the response becomes a Python warning, not an exception ([api_extractor.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/api_extractor.py#L18-L51)). It uses blocking `requests` inside an `async` method, so `arun` calls on this path block the event loop.

## Extending it

- **Any LangChain model.** Pass `model=` with an OpenAI, Azure, Ollama or other `BaseChatModel`. `HuggingFaceModel` wraps a local `transformers` pipeline and runs it in a thread executor ([model.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/model.py#L47-L112)). Note that the bundled wrappers are `@singleton`s, so the first constructor arguments win for the whole process.
- **Custom extractor.** Subclass `Extractor` (one async `run(content, attributes, prompt)`) or a `LocalExtractor` with your own `system_prompt`, and pass it as `extractor=`. `ChunksTabularExtractor` also accepts `chunk_size`, a `token_counter` and a custom `MarkdownConverter`.
- **Playwright hooks.** `initial_script` runs once per session, and `playwright_script` runs after every navigation. Both receive and return a `Page`, so clicks, logins and waits go there.
- **Direct loader use.** `PageLoader(browser=...)` accepts your own Playwright browser, for example a headed one with `slow_mo` for debugging, as the bundled scrolling example does.

## Running it

- `pip install parsera` and `playwright install` (Firefox is the browser that is used). Python 3.10 or newer.
- Hosted mode: set `PARSERA_API_KEY` and call `Parsera().run(url=..., elements=...)`.
- Local mode: pass `model=`, and set whatever key that provider needs.
- CLI: `python -m parsera.main URL --scheme '{"title":"h1"}' --scrolls 5 --output out.json`. The CLI always builds `GPT4oMiniModel`, so it needs `OPENAI_API_KEY`, not a Parsera key ([main.py](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/main.py#L121-L135)).
- Docker: the image installs Poetry and Playwright Firefox and uses the CLI as its entrypoint. `docker-compose.yaml` mounts a `scheme.json` and an `output/` folder.

## Strengths and caveats

- **Strength: tiny and readable.** The whole data path fits in five files. It is easy to audit, fork, or lift the chunk-and-merge idea from.
- **Strength: chunking that respects rows.** Overlapping chunks, "continue this sequence" prompts and a merge call are a reasonable answer to tables longer than a context window.
- **Strength: lazy-load capture.** The removed-node observer recovers rows that virtualised lists delete while scrolling, which many scrapers silently miss.
- **Caveat: the default is a paid cloud call.** `Parsera()` with no model sends the whole page HTML to a third-party API. You must pass `model=` to keep data local.
- **Caveat: full LLM cost on every page.** Nothing is cached or turned into selectors. A large page that splits into N chunks costs N + 1 sequential calls, and the merge call re-sends all extracted data.
- **Caveat: weak anti-bot.** Firefox plus a Chrome-oriented UA patch plus `playwright-stealth` is basic. There are no CAPTCHA, retry or backoff features.
- **Caveat: session reuse is fragile.** `run()` calls `asyncio.run` each time, while the Playwright session from the first call is kept on the loader. Repeated sync calls on one instance can hit a session bound to a closed event loop. Use `arun` for multi-page jobs, or a fresh instance per URL.
- **Caveat: hygiene.** The committed `run.py` sets a literal `PARSERA_API_KEY` value, and the test suite needs live Azure OpenAI credentials and network access.

*Sources: code at 865d777, deepwiki-open wiki (10 pages), OpenDeepWiki wiki (9 pages), verified Q&A.*

## How raznem/parsera answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

Pages are fetched exclusively via **Playwright** controlling a headless **Firefox** browser — there is no plain-HTTP fallback. The `PageLoader` class (`parsera/page.py:19`) manages the Playwright lifecycle: it launches Firefox headless (`page.py:37`), creates a browser context, opens a page, navigates to the URL with `page.goto()`, and waits for a load state (default `"networkidle"`, also accepts `"domcontentloaded"` or `"load"`, with a try/except pass on timeout at `page.py:183-190`). JavaScript rendering is therefore fully supported — all JS on the page executes before content is captured. For JS-heavy or infinite-scroll pages, the `scroll_page()` method (`page.py:92-144`) scrolls programmatically with configurable scroll-limit, a 2-second sleep between scrolls, and a MutationObserver that captures elements removed during lazy-loading. Content is retrieved as raw HTML via `document.documentElement.outerHTML`, including content from all iframes fetched in parallel (`page.py:155-173`). There is no support for PDF or image content — the library reads HTML only. A custom Playwright script can be injected at session creation (`initial_script`) or before each page fetch (`playwright_script`), enabling auth flows, cookie injection, or custom wait logic (`page.py:78-90`, `parsera.py:60-71`).


Citations: [parsera/page.py:29-37](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L29-L37) · [parsera/page.py:155-201](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L155-L201) · [parsera/page.py:92-144](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L92-L144) · [parsera/parsera.py:54-77](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/parsera.py#L54-L77)

### How is content extracted or converted? (answered)

Extraction follows a two-stage pipeline: HTML→Markdown conversion via **markdownify** (`MarkdownConverter`), then LLM-based structured parsing. Every extractor in `parsera/engine/simple_extractor.py` inherits from `LocalExtractor` and calls `self.converter.convert(content)` before sending to the LLM (`simple_extractor.py:68`). There are no readability-style heuristics or boilerplate-removal rules — the full page markdown is sent to the model, which is expected to ignore irrelevant content. Five extractor types exist: **TabularExtractor** returns a list-of-dicts (`[{...}, {...}]`) for table-like data (`simple_extractor.py:156-157`). **ListExtractor** returns dict-of-lists (`simple_extractor.py:202-203`). **ItemExtractor** returns a single dict (`simple_extractor.py:248-249`). **ChunksTabularExtractor** (`chunks_extractor.py:184`) extends TabularExtractor with recursive character text-splitting (LangChain's `RecursiveCharacterTextSplitter` with ~33% overlap, `chunk_size` default 100k tokens) for pages that exceed context windows; it extracts each chunk sequentially, passing previous results as context to handle truncated rows, then merges all chunk results via a second LLM call (`chunks_extractor.py:274-292`). **StructuredExtractor** (`structured_extractor.py:31`) extends ChunksTabularExtractor with Pydantic schema-based structured output: it dynamically creates a `pydantic.BaseModel` from user-supplied `type`/`description` fields and calls `model.with_structured_output()` (`structured_extractor.py:102-142`). The user specifies fields as a dict of `field_name → description` (Tabular) or `field_name → {type, description}` (Structured).


Citations: [parsera/engine/simple_extractor.py:39-86](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/simple_extractor.py#L39-L86) · [parsera/engine/chunks_extractor.py:184-331](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/chunks_extractor.py#L184-L331) · [parsera/engine/structured_extractor.py:31-146](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/structured_extractor.py#L31-L146) · [parsera/engine/simple_extractor.py:156-249](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/simple_extractor.py#L156-L249)

### How are LLMs used, if at all? (answered)

LLMs are the core extraction engine — every extractor sends the full page markdown (or chunks thereof) to a **LangChain BaseChatModel** (`langchain_core.language_models`) via `await model.ainvoke(messages)`, using a `SystemMessage` (the extractor's behaviour prompt with few-shot examples) + `HumanMessage` (the markdown content and field definitions). The default model when none is provided is **GPT-4o-mini** (`model.py:36-44`), wrapped as a singleton `GPT4oMiniModel`. Alternatively users can pass any LangChain-compatible chat model: **Azure OpenAI** via `AzureModel` (`model.py:20-32`), **Ollama** models (documented in `docs/features/custom-models.md`), or **HuggingFace pipelines** via `HuggingFaceModel` (`model.py:48-112`), which runs the pipeline in a thread executor since it lacks native async. Chunking is used for large pages: `ChunksTabularExtractor` uses `RecursiveCharacterTextSplitter` with a default `chunk_size=100000` tokens (`chunks_extractor.py:227-231`), counted via `tiktoken` with the `o200k_base` encoding. When chunked, each chunk first extends/continues the previous extraction via an "append" prompt, then a final LLM call merges all chunk results deduplicating by proximity to boundaries. For structured output, the `StructuredExtractor` calls `model.with_structured_output(schema=OutputSchema)` on a dynamically-generated Pydantic model (`structured_extractor.py:142`). There is no explicit cost control mechanism beyond the user choosing a smaller model or limiting chunk size. The `APIExtractor` (`api_extractor.py:18-51`) bypasses local LLM usage entirely, sending the raw HTML content to `https://api.parsera.org/v1/parse` with a `PARSERA_API_KEY`, making the cloud service the LLM backend.

> **Editor's note.** Correction: with no `model` and no `extractor`, `Parsera()` uses `APIExtractor`, which sends the raw page HTML to the hosted api.parsera.org endpoint with `PARSERA_API_KEY` (parsera/parsera.py L36-L37). `GPT4oMiniModel` is only the default of the CLI in parsera/main.py; library users must pass `model=` to run extraction locally.

Citations: [parsera/engine/model.py:36-44](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/model.py#L36-L44) · [parsera/engine/chunks_extractor.py:211-231](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/chunks_extractor.py#L211-L231) · [parsera/engine/chunks_extractor.py:294-331](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/chunks_extractor.py#L294-L331) · [parsera/engine/structured_extractor.py:102-146](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/structured_extractor.py#L102-L146) · [parsera/engine/api_extractor.py:18-51](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/api_extractor.py#L18-L51)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

Anti-bot evasion is handled through a dedicated **stealth mode** (enabled by default) using the `playwright-stealth` library. In `PageLoader.stealth()` (`page.py:40-66`), after the initial context is created and a page opened, the library replaces the User-Agent string to remove the `HeadlessChrome/` marker — it evaluates `navigator.userAgent`, strips `HeadlessChrome/` down to `Chrome/` (`page.py:44-45`), then closes the original context and creates a new one with the spoofed User-Agent. It then applies `stealth_async()` with a `StealthConfig` that skips overriding `navigator_user_agent`, `navigator_plugins`, and `navigator_vendor` (since the UA is already patched manually, `page.py:57-63`). **Proxy support** is built-in: a `ProxySettings` TypedDict (`server`, optional `bypass`/`username`/`password`) can be passed to `run()`/`arun()`, and when creating the Playwright context, `proxy=proxy_settings` is forwarded to Playwright's `new_context()` (`page.py:48-49`, `page.py:76`). **CAPTCHA handling** is not implemented in the codebase. **Rate limiting** is not implemented — there is no throttle, backoff, or concurrency limiter on page fetches. **Cookie injection** is supported: custom cookies passed to the constructor (`custom_cookies`) are added to the browser context via `context.add_cookies()` (`page.py:51-55`, `page.py:78-82`) with validation that raises `CookiesValidationException` on failure. The browser is always **headless** Firefox (`page.py:37`), though examples show `headless=False` is possible with a custom `PageLoader` or by passing a custom browser.


Citations: [parsera/page.py:40-66](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L40-L66) · [parsera/page.py:68-90](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L68-L90) · [parsera/parsera.py:36-52](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/parsera.py#L36-L52) · [parsera/page.py:12-17](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/page.py#L12-L17)

### How is crawling at scale implemented? (not applicable)

Parsera does not implement crawling at scale. It is a single-page scraping library with no queue system, no URL deduplication, no depth or crawl-limit parameters, no robots.txt parsing, no politeness delays, and no distributed worker support. Each `run()`/`arun()` call fetches exactly one URL, extracts its data, and returns — the user must build any multi-page or crawl logic externally. There is no concurrency model for page fetching (Playwright handles one page at a time per session). The only "scale" feature is page-level chunking for long content, which addresses LLM context limits, not crawling breadth.


Citations: [parsera/parsera.py:79-97](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/parsera.py#L79-L97) · [parsera/parsera.py:99-115](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/parsera.py#L99-L115)

### What is the developer interface? (answered)

The primary interface is the **Python library API** via the `Parsera` class (`parsera/parsera.py:13-115`). Instantiate with optional `model` (LangChain chat model), `extractor`, `initial_script` (Playwright callback), `stealth` (bool), and `custom_cookies`. Call `.run(url, elements, prompt, ...)` for synchronous usage (wraps `asyncio.run()`) or `.arun(...)` for async. The result is a `list[dict]` (for tabular extractors) or `dict` (for Item/List extractors). A **CLI** is provided via `python -m parsera.main` (`main.py:56-181`), accepting a URL, `--scheme` (JSON string of field definitions), `--file` (JSON file with same), `--scrolls` (int), and `--output` (file path, default "output.json"). The CLI uses `GPT4oMiniModel` by default and requires `OPENAI_API_KEY` or `PARSERA_API_KEY` in the environment. A **Docker** entrypoint is configured (`Dockerfile:42`) that runs the CLI with environment variables for URL, FILE, and OUTPUT. There is **no REST API service** served by the project (the `APIExtractor` is a *client* to a hosted API at `api.parsera.org`, not a server). There is **no MCP server**, **no web UI**, and **no graphical interface**. Output formats are limited to Python dicts (in-code) and JSON files (CLI). Language bindings are **Python-only**. The project exposes no language-agnostic protocol.

> **Editor's note.** Correction: the CLI always constructs `GPT4oMiniModel` (parsera/main.py L121-L135), so it needs `OPENAI_API_KEY`; a `PARSERA_API_KEY` alone does not make the CLI work.

Citations: [parsera/parsera.py:13-97](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/parsera.py#L13-L97) · [parsera/main.py:56-180](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/main.py#L56-L180) · [parsera/engine/api_extractor.py:18-51](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/parsera/engine/api_extractor.py#L18-L51) · [Dockerfile:1-42](https://github.com/raznem/parsera/blob/865d7772fc92af9dfe3622ca1adcd5e5b7afd07f/Dockerfile#L1-L42)
