# ScrapeGraphAI/Scrapegraph-ai

> Python library that turns a prompt and a URL into JSON by running LangChain LLMs over fetched, chunked pages in node graphs.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/ScrapeGraphAI/Scrapegraph-ai (reviewed at commit `194055e203afce41ed4e70365dbc416bad756115`, 2026-10-06)
- Stars: 31573 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/scrapegraph-ai/

## Overview

ScrapeGraphAI is a Python library that writes no scraping logic by hand: you give it a natural-language prompt, a source (URL, local file or raw text) and an LLM config, and it returns JSON. Each task is a "graph", a small directed pipeline of nodes such as fetch, parse and generate-answer. `SmartScraperGraph` handles one page. `SearchGraph` searches the web first. `SmartScraperMultiGraph` fans out over URLs. There are also variants for CSV, JSON and XML, depth-limited crawling, screenshots, speech and script generation.

The engineering weight is in prompting and orchestration, not fetching. Pages are loaded with Playwright (plus a stealth patch), converted to text or Markdown, cut into token-sized chunks, and sent to any LangChain chat model. With several chunks, each chunk is answered in parallel and a final "merge" prompt combines the partial answers. Without a schema, output is free-form JSON. With a Pydantic schema, the model gets format instructions and the reply is parsed against them.

It is a good fit for quick prompt-driven extraction from a handful of pages. It is not a crawler framework. There is no persistent queue, no robots.txt enforcement in the built-in graphs, and no retry or proxy strategy beyond what you configure on the browser loader.

## Architecture

```mermaid
flowchart LR
  U["SmartScraperGraph(prompt, source, config)"] --> AG["AbstractGraph: _create_llm"]
  AG --> BG["BaseGraph.execute"]
  BG --> FN["FetchNode"]
  FN --> CL["ChromiumLoader (Playwright)"]
  FN --> ALT["BrowserBase / ScrapeDo / Plasmate"]
  FN --> PN["ParseNode: html2text + semchunk"]
  PN --> GA["GenerateAnswerNode"]
  GA --> LLM["LangChain chat model"]
  GA --> MG["merge prompt"]
  BG --> TL["telemetry (on by default)"]
  SG["SearchGraph"] --> SI["SearchInternetNode"]
  SI --> GI["GraphIteratorNode"]
  GI --> U
```

| Component | Path | Role |
|---|---|---|
| Graph base | `scrapegraphai/graphs/abstract_graph.py` | Builds the LLM from config, sets token budget, pushes common params to nodes |
| Executor | `scrapegraphai/graphs/base_graph.py` | Walks nodes along edges, tracks token cost per node, sends telemetry |
| Graphs | `scrapegraphai/graphs/*.py` | About 25 ready pipelines (`SmartScraperGraph`, `SearchGraph`, `DepthSearchGraph`, ...) |
| Nodes | `scrapegraphai/nodes/` | Fetch, parse, generate-answer, conditional, search, iterator, merge, code generation |
| Loaders | `scrapegraphai/docloaders/` | `ChromiumLoader` (Playwright or Selenium), BrowserBase, ScrapeDo, Plasmate |
| Prompts | `scrapegraphai/prompts/` | Templates for single-chunk, per-chunk and merge calls (HTML and Markdown variants) |
| Utils | `scrapegraphai/utils/` | `convert_to_md`, `cleanup_html`, `split_text_into_chunks`, output parsers, proxy broker, web search |
| Models | `scrapegraphai/models/`, `helpers/models_tokens.py` | Wrappers for non-LangChain providers; context-window table per model |
| Telemetry | `scrapegraphai/telemetry/telemetry.py` | Posts run data to the maintainers' tracing endpoint |

## How a request flows

Take `SmartScraperGraph("List the products", "https://example.com", {"llm": {"model": "openai/gpt-4o-mini"}}).run()`:

1. **Build the LLM.** `AbstractGraph.__init__` calls `_create_llm`. It splits `provider/model`, or guesses the provider from the token table, checks it against a fixed allow-list, and sets `model_token` from `models_tokens` (falling back to 8192 with a warning). It then calls LangChain's `init_chat_model` or a custom wrapper for DeepSeek, xAI, Nvidia and others ([abstract_graph.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/abstract_graph.py#L56-L292)).
2. **Pick a graph shape.** `_create_graph` chooses one of eight node layouts from the `html_mode`, `reasoning` and `reattempt` flags. The default is `FetchNode -> ParseNode -> GenerateAnswerNode` ([smart_scraper_graph.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/smart_scraper_graph.py#L58-L297)).
3. **Execute.** `BaseGraph._execute_standard` starts at the entry node and runs each node inside a token-counting callback. It follows `edges`, or the node name a `ConditionalNode` returns ([base_graph.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/base_graph.py#L198-L340)).
4. **Fetch.** `FetchNode.handle_web_source` uses BrowserBase, ScrapeDo or Plasmate if configured. Otherwise it uses `ChromiumLoader`. For OpenAI and Azure models (or with `force`), the HTML is converted to Markdown with `html2text` ([fetch_node.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/fetch_node.py#L266-L409)).
5. **Load the page.** `ChromiumLoader.ascrape_playwright` launches Chromium (or Firefox) with an optional proxy, applies `Malenia.apply_stealth` from `undetected-playwright` to the context, goes to the URL with `domcontentloaded`, waits for `load_state` and returns `page.content()`, retrying up to `retry_limit` (default 1) ([chromium.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/docloaders/chromium.py#L349-L410)).
6. **Parse and chunk.** `ParseNode` runs LangChain's `Html2TextTransformer` (links kept) and splits the text with `semchunk` to `model_token - 250` tokens. It also warns when none of the prompt's or schema's terms appear in the text ([parse_node.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/parse_node.py#L127-L220)).
7. **Answer.** `GenerateAnswerNode` builds format instructions from the schema, or asks for `{"content": ...}`. One chunk means one call. Several chunks become a `RunnableParallel` of per-chunk chains, then one merge call over all partial answers ([generate_answer_node.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/generate_answer_node.py#L120-L269)). `run()` returns `state["answer"]`.

## Key components

### Chunking and merging

`split_text_into_chunks` uses `semchunk` with a token counter and a further 10% safety margin. A word-based splitter is kept as a fallback ([split_text_into_chunks.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/utils/split_text_into_chunks.py#L10-L59)). The map-then-merge pattern scales to long pages. But each chunk's answer is produced without seeing the others, and the merge step is one more unconstrained LLM call that can drop or invent rows.

### Structured output

Schema support is prompt-level. `get_pydantic_output_parser` returns a LangChain `JsonOutputParser` for Pydantic v2 models and rejects v1 ([output_parser.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/utils/output_parser.py#L70-L92)). The schema becomes format instructions in the prompt, and the reply is parsed as JSON. Provider-side constrained decoding is used only for Ollama, where the node sets `llm_model.format` to the schema's JSON Schema ([generate_answer_node.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/generate_answer_node.py#L60-L72)). Bedrock gets no format instructions at all.

### Multi-page graphs

`SearchGraph` chains `SearchInternetNode` (DuckDuckGo by default, or Bing, SearXNG or Serper), `GraphIteratorNode` and `MergeAnswersNode` ([search_graph.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/search_graph.py#L55-L100)). The iterator creates one `SmartScraperGraph` per URL and runs each `graph.run` in a thread under an `asyncio.Semaphore` (default batch size 16) ([graph_iterator_node.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/graph_iterator_node.py#L95-L147)). `DepthSearchGraph` uses `FetchNodeLevelK`, which repeats link extraction for `depth` rounds over an in-memory document list ([fetch_node_level_k.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/fetch_node_level_k.py#L72-L102)).

### robots.txt

`RobotsNode` fetches `/robots.txt` and asks the LLM, with a prompt, whether the path is allowed for an agent name looked up in `helpers/robots.py`. It raises unless `force_scraping` is set ([robots_node.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/robots_node.py#L80-L131)). None of the shipped graphs include it, so in practice nothing checks robots.txt unless you build a custom graph.

### Telemetry

After every successful run, `BaseGraph` calls `log_graph_execution` with the prompt, schema, parsed page content, LLM answer, model name and URL ([base_graph.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/base_graph.py#L310-L342)). Telemetry is on by default. When all fields are present, which is any URL run with a schema, the payload is posted to an external tracing endpoint, capped at 1000 calls per session ([telemetry.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/telemetry/telemetry.py#L77-L145), [L177-L215](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/telemetry/telemetry.py#L177-L215)). Set `SCRAPEGRAPHAI_TELEMETRY_ENABLED=false` if scraped content or prompts are sensitive.

## Extending it

- **Custom graphs.** Subclass `AbstractGraph`, implement `_create_graph` to return a `BaseGraph(nodes, edges, entry_point)`, and `run`. Nodes declare inputs with a small boolean expression language (`"user_prompt & (relevant_chunks | parsed_doc | doc)"`) that picks keys from the shared state dict.
- **Custom nodes.** Subclass `BaseNode` and implement `execute(state) -> state`.
- **Any model.** Pass `{"model_instance": your_langchain_chat_model, "model_tokens": N}` to skip the provider allow-list ([abstract_graph.py](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/abstract_graph.py#L150-L156)).
- **Fetch backends.** `loader_kwargs` reach `ChromiumLoader` (`backend="selenium"`, `requires_js_support`, `load_state`, `proxy`, `retry_limit`). `browser_base`, `scrape_do` and `plasmate` switch to hosted fetchers.
- **Burr.** `burr_kwargs` runs the graph through the Burr state-machine framework for tracking and replay.

## Running it

- `pip install scrapegraphai`, then `playwright install` for the browser. Python 3.12 or newer is required.
- An LLM is mandatory: an API key for a hosted provider, or a local Ollama model (`"model": "ollama/llama3.2"`). Set `model_tokens` for models missing from the table, or long pages will be chunked against an 8192-token default.
- There is no server, CLI or MCP endpoint in this repository. Everything runs in-process. The `docker-compose.yml` only bundles an Ollama container.

## Strengths and caveats

- **Strength: very little code per task.** A prompt and an optional Pydantic model replace selectors, and the graph library covers common shapes (single page, search, multi-URL, files).
- **Strength: provider-agnostic.** Anything LangChain can instantiate works, including local models through Ollama.
- **Strength: clear pipeline model.** Nodes and a shared state dict make it easy to add a reasoning step, a retry branch or a custom node.
- **Caveat: an LLM call on every page.** Cost and latency scale with page size and chunk count. Nothing generates reusable selectors for the default scrape path.
- **Caveat: telemetry ships content off-box by default.** Prompts, page text and answers go to the maintainers' endpoint unless you disable it.
- **Caveat: thin fetching.** One retry attempt by default, a basic stealth patch, no block detection and no robots.txt in the shipped graphs. The `use_soup` HTTP path in `FetchNode` also references variables it never sets, so it fails if enabled.
- **Caveat: dead cloud shortcut.** `SmartScraperGraph` compares the model object to the string `"scrapegraphai/smart-scraper"` to forward to the hosted API. That branch can never run, because `_create_llm` rejects the `scrapegraphai` provider first.

*Sources: code at 194055e, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (12 pages), verified Q&A.*

## How ScrapeGraphAI/Scrapegraph-ai answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

**Plain HTTP vs. headless browser.** `FetchNode` has two paths controlled by the `use_soup` config flag (default `False`). When `use_soup=True`, it uses `requests.get` (plain HTTP, no JS). When `use_soup=False`, it delegates to `ChromiumLoader`, which spins up a real headless Chromium via Playwright (or Selenium/undetected-chromedriver with the `backend` parameter). 

**JS rendering.** The ChromiumLoader has two scraping modes: `ascrape_playwright` waits for `"domcontentloaded"`, while `ascrape_with_js_support` waits until `"networkidle"`, ensuring JavaScript-rendered content is fully loaded. The `FetchNode` also passes `requires_js_support` config to control this. There is also an `ascrape_playwright_scroll` variant that uses mouse-wheel events and configurable sleep intervals for lazy-loaded/infinite-scroll pages.

**Waiting strategy.** The `load_state` parameter (default `"domcontentloaded"`) controls how long Playwright waits before capturing page content. The node also supports `retry_limit` (attemps) and `timeout` (blocking-ops seconds).

**Supported content types.** Beyond HTML, `FetchNode.handle_file` loads PDFs (via `PyPDFLoader`), CSVs (pandas), JSON, XML, and Markdown files from local paths. It also handles entire directories of these types, and local HTML strings via `handle_local_source`. Search-engine results from `SearchInternetNode` return plain lists of URLs.

**Third-party fetching services.** The node supports `browser_base` (BrowserBase API), `scrape_do` (ScrapeDo proxy API), and `plasmate` (PlasmateLoader) as alternative fetching backends.


Citations: [scrapegraphai/nodes/fetch_node.py:93-126](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/fetch_node.py#L93-L126) · [scrapegraphai/nodes/fetch_node.py:266-409](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/fetch_node.py#L266-L409) · [scrapegraphai/docloaders/chromium.py:349-406](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/docloaders/chromium.py#L349-L406) · [scrapegraphai/docloaders/chromium.py:409-464](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/docloaders/chromium.py#L409-L464) · [scrapegraphai/docloaders/chromium.py:191-347](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/docloaders/chromium.py#L191-L347)

### How is content extracted or converted? (answered)

**HTML→Markdown conversion.** `ParseNode.execute` uses LangChain's `Html2TextTransformer` to convert HTML to plain text with links preserved. Separately, `FetchNode` conditionally calls `convert_to_md()` which uses the `html2text` library to produce Markdown from raw HTML, enabling either full-Markdown or text-only downstream processing. The `cut=True` flag controls whether the cleanup step is applied.

**Boilerplate removal / HTML cleanup.** The `cleanup_html()` function (BeautifulSoup-based) extracts the `<title>`, removes `<style>` tags, minifies body HTML via the `minify` library, extracts JSON from `<script>` tags, and collects link/image URLs. A separate `reduce_html()` function applies increasing levels of reduction: level 0 = regex minification, level 1 = remove comments and non-essential attributes, level 2 = truncate text content to 20 chars per tag. Default filters in `helpers/default_filters.py` provide common image-extension lists.

**Content-chunking with schema awareness.** After HTML→text conversion, `ParseNode` calls `split_text_into_chunks()` which uses the `semchunk` library (or word-level fallback) to split content into token-budgeted chunks sized to the LLM's `model_token` limit. The chunk_size is reduced by 250 (HTML mode) or 500 (other modes) to leave room for prompt instructions. The parser also collects URL fields that contain `[link URLs, image URLs]` for schema-aware follow-up.

**Schema-based extraction.** Extraction shaping is done downstream in `GenerateAnswerNode`. If the user provides a Pydantic `schema`, the LLM receives format instructions via `get_pydantic_output_parser`, which forces structured JSON output matching the schema. Without a schema, a `TolerantJsonOutputParser` is used with a default JSON format instruction.

**Content-evidence warning.** `ParseNode` has a `_warn_if_content_lacks_requested_fields` method that deterministically checks whether the user's prompt terms or schema field names appear in the parsed text — flagging empty/error-page/dropped-content conditions before they reach the LLM.


Citations: [scrapegraphai/nodes/parse_node.py:150-219](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/parse_node.py#L150-L219) · [scrapegraphai/utils/split_text_into_chunks.py:10-59](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/utils/split_text_into_chunks.py#L10-L59) · [scrapegraphai/nodes/generate_answer_node.py:120-175](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/generate_answer_node.py#L120-L175)

### How are LLMs used, if at all? (answered)

**Extensive LLM usage across every pipeline stage.** The library uses LLMs not just for final answer generation but also for search-query formulation (`SearchInternetNode`), HTML description (`DescriptionNode`), code generation (`GenerateCodeNode`), prompt refinement (`PromptRefinerNode`), URL relevance ranking (`SearchLinkNode`), and decision routing (`ConditionalNode`). Every `GenerateAnswerNode` chains a LangChain `PromptTemplate` to an LLM.

**Prompting.** All prompts come from `scrapegraphai/prompts/` — distinct templates exist for single-chunk (`TEMPLATE_NO_CHUNKS`/`TEMPLATE_NO_CHUNKS_MD`), multi-chunk (`TEMPLATE_CHUNKS`/`TEMPLATE_CHUNKS_MD`), and merge (`TEMPLATE_MERGE`/`TEMPLATE_MERGE_MD`) workflows. Additional info can be prepended. When `is_md_scraper` or `not script_creator` is true, the Markdown variants (which ask the LLM to reason over Markdown-syntax content) are used.

**Chunking large pages.** `GenerateAnswerNode.execute` checks `len(doc)`: a single chunk goes directly to the LLM; multiple chunks are processed concurrently via `RunnableParallel` (one chain per chunk), then a final merge prompt sends all partial results to the LLM for synthesis. Chunk size is driven by the model's max-token cap from `helpers/models_tokens.py`.

**Structured output / JSON schema.** When a Pydantic `schema` is provided: for `ChatOpenAI` models it uses `get_pydantic_output_parser` (which binds `response_format` or function-calling). For `ChatOllama`, it sets `llm_model.format` to the schema's JSON schema. For `ChatBedrock`, no format instructions are passed. Without a schema, `TolerantJsonOutputParser` wraps the response in `{"content": ...}`.

**Supported providers.** The `_create_llm` method in `AbstractGraph` supports 20+ providers: OpenAI, Azure, Google Gemini/Vertex, Ollama, Anthropic, Bedrock, MistralAI, Groq, Hugging Face, DeepSeek, TogetherAI, Fireworks, Ernie, plus custom models (CLoD, DeepSeek, MiniMax, Nvidia, OneApi, XAI). Provider is auto-detected from model name or specified via `provider/model` format.

**Cost controls.** `AbstractGraph` supports `InMemoryRateLimiter` (requests per second) and `max_retries` for rate limiting. Costs are tracked per-model in `utils/model_costs.py` (input cost per 1K tokens). The `BaseGraph` execution info reports total tokens, prompt/completion tokens, and cumulative USD cost per run. The `model_tokens` config parameter lets users override the default window (falls back to 8192 with a warning).

> **Editor's note.** Correction: get_pydantic_output_parser returns a plain LangChain JsonOutputParser for every non-Bedrock model; it does not bind response_format or function calling. The schema only becomes format instructions in the prompt, and only Ollama gets provider-side constraint (llm_model.format set to the JSON Schema).

Citations: [scrapegraphai/graphs/abstract_graph.py:121-291](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/abstract_graph.py#L121-L291) · [scrapegraphai/nodes/generate_answer_node.py:120-269](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/generate_answer_node.py#L120-L269) · [scrapegraphai/helpers/models_tokens.py:5-73](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/helpers/models_tokens.py#L5-L73) · [scrapegraphai/utils/model_costs.py:1-50](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/utils/model_costs.py#L1-L50) · [scrapegraphai/graphs/base_graph.py:198-220](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/base_graph.py#L198-L220)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

**Stealth patches.** The ChromiumLoader uses `undetected-playwright`'s `Malenia.apply_stealth()` on every browser context (`chromium.py:392`). This patches Playwright's browser fingerprint to avoid bot detection. For Selenium, it uses `undetected_chromedriver` (uc) which patches the ChromeDriver to bypass anti-bot systems.

**User-agent rotation.** `research_web.py` maintains a hardcoded list of 5 real-looking user-agent strings (Chrome, Safari, Firefox, mobile variants) and selects one randomly via `get_random_user_agent()` for search-engine HTTP requests. This is separate from the browser backend.

**Proxy rotation.** The `utils/proxy_rotation.py` module provides full proxy brokering: `search_proxy_servers()` uses the `free-proxy` library (`FreeProxy`) to discover proxies matching criteria (anonymity, country set, HTTPS support, timeout). It tests each candidate by making a request and returns verified working proxies. The `parse_or_search_proxy()` function routes to either a known proxy server or the broker. Proxies are passed to Playwright's browser `launch()` method as the `proxy` parameter.

**Third-party proxy services.** `FetchNode` supports proxy-gated fetching via `scrape_do` (ScrapeDo API with geoCode/superProxy options) and `browser_base` (BrowserBase cloud browser API). `PlasmateLoader` is another alternative backend with its own header/proxy settings.

**CAPTCHA handling.** There is no explicit CAPTCHA solving module in the codebase. The combination of undetected-playwright stealth patches and undetected-chromedriver is the primary defense against CAPTCHA triggers. `storage_state` can persist authenticated browser sessions to avoid re-auth.

**Rate limiting.** `research_web.py` has a `@rate_limited` decorator (default: 10 calls per 60 seconds) applied to `search_on_web` for polite search-engine requests. LLM calls can be rate-limited via LangChain's `InMemoryRateLimiter`.

**HTTP error detection.** Both the `use_soup` path (response.status_code check) and ChromiumLoader (`_warn_on_error_status`) detect 4xx/5xx responses and log warnings, preventing error-page content from silently reaching the LLM.


Citations: [scrapegraphai/docloaders/chromium.py:363-397](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/docloaders/chromium.py#L363-L397) · [scrapegraphai/utils/proxy_rotation.py:47-131](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/utils/proxy_rotation.py#L47-L131) · [scrapegraphai/utils/proxy_rotation.py:191-216](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/utils/proxy_rotation.py#L191-L216) · [scrapegraphai/utils/research_web.py:92-119](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/utils/research_web.py#L92-L119) · [scrapegraphai/utils/research_web.py:140-157](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/utils/research_web.py#L140-L157)

### How is crawling at scale implemented? (answered)

**Link-level crawling (depth-limited).** The `DepthSearchGraph` pipeline uses `FetchNodeLevelK`, which implements recursive BFS-style crawling: it fetches a URL, extracts all `<a href=...>` links via BeautifulSoup, resolves relative URLs, deduplicates against already-seen URLs, then fetches the new links. This repeats for `depth` iterations. The `only_inside_links` flag restricts to same-domain links. The `FetchNodeLevelK.obtain_content()` method manages the frontier as a document list where each entry is `{"source": url, "document": ...}`.

**No robots.txt enforcement.** Despite a `helpers/robots.py` file, it only maps model names to bot user-agent strings (e.g., `"gpt-4o"`→`"GPTBot"`). There is no actual robots.txt fetching, parsing, or politeness-delay implementation in the code. A `robots_node.py` module exists but bundles an LLM prompt rather than implementing `robots.txt` parsing.

**Concurrency via asyncio semaphore.** The `GraphIteratorNode` runs multiple graph instances concurrently using `asyncio.Semaphore(batchsize)` (default 16). For each URL in the input list, it spawns a `graph.run()` call in a thread via `asyncio.to_thread`, and limits parallel execution with the semaphore. Tqdm progress bars track completion.

**URL deduplication.** In `FetchNodeLevelK.obtain_content()`, newly extracted links are checked against both the `documents` and `new_documents` lists before being added to the frontier: `if not any(d.get("source") == link for d in documents) and not any(d.get("source") == link for d in new_documents)`. There is no shared visited set across graph runs.

**No distributed workers.** The codebase has no message-queue, Redis, or distributed-worker infrastructure. All concurrency is in-process via asyncio. The `batch_api.py` utility hints at batched LLM API calls but does not distribute scraping work.

**Max results limits.** Graphs accept `max_results` (default 3 in `SearchGraph`) and `max_results` in `SearchInternetNode` to cap the number of fetched pages. The `AbstractGraph` config can set `max_results` for overall graph iteration counts in graph memory.

> **Editor's note.** Correction: RobotsNode does fetch /robots.txt and asks the LLM whether the path is allowed (raising unless force_scraping), but no shipped graph includes it, so default runs never consult robots.txt.

Citations: [scrapegraphai/nodes/fetch_node_level_k.py:72-102](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/fetch_node_level_k.py#L72-L102) · [scrapegraphai/nodes/fetch_node_level_k.py:234-273](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/fetch_node_level_k.py#L234-L273) · [scrapegraphai/nodes/graph_iterator_node.py:46-147](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/nodes/graph_iterator_node.py#L46-L147) · [scrapegraphai/helpers/robots.py:1-14](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/helpers/robots.py#L1-L14) · [scrapegraphai/graphs/depth_search_graph.py:75-87](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/depth_search_graph.py#L75-L87)

### What is the developer interface? (answered)

**Library API (primary).** The main developer interface is importing graph classes from `scrapegraphai.graphs` and running them: `SmartScraperGraph(prompt, source, config).run()`. The 25+ graph variants handle different use cases — single-page (`SmartScraperGraph`), multi-page (`SmartScraperMultiGraph`), search-engine (`SearchGraph`), depth-recursive (`DepthSearchGraph`), CSV/JSON/XML (`CSVScraperGraph`, `JSONScraperGraph`, `XMLScraperGraph`), screenshot (`ScreenshotScraperGraph`), image-to-text (`OmniScraperGraph`), speech output (`SpeechGraph`), and code-generation (`ScriptCreatorGraph`). Config is a plain dict with `llm`, `headless`, `verbose`, etc.

**Graph builder.** The `builders/graph_builder.py` provides `GraphBuilder` for constructing custom pipelines from YAML configs without writing Python code.

**Burr integration.** The `integrations/burr_bridge.py` allows executing graphs via the Burr orchestration framework when `burr_kwargs` is passed in config, providing state-persistence and replay capabilities.

**CLI.** There is no dedicated CLI entry point in `pyproject.toml` or setup files. The library is used programmatically. However, the integrations include a `scrapegraph_py_compat` module that proxies `SmartScraperGraph` to the cloud ScrapeGraphAI service when `llm_model == "scrapegraphai/smart-scraper"`.

**Output formats.** Results are Python dicts (JSON). The `utils/data_export.py` supports CSV/JSON export. `utils/save_audio_from_bytes.py` handles audio output for speech graphs. `utils/save_code_to_file.py` saves generated scraping scripts.

**Language bindings.** Python only. The `pyproject.toml` requires Python ≥3.12. No REST service, MCP server, or UI is included in the open-source repository. The `scrapegraph-py` SDK dependency (≥2.0.0) connects to https://scrapegraphai.com cloud API when configured, but the open-source code runs locally via LangChain and Playwright.

**Error reporting.** The execution info dict returned by `BaseGraph` includes per-node token counts, execution time, and cumulative USD cost, accessible via `graph.get_execution_info()`.

> **Editor's note.** Correction: the scrapegraph_py_compat cloud shortcut in SmartScraperGraph is unreachable: _create_llm rejects the 'scrapegraphai' provider first, and llm_model is a model object that never equals the string 'scrapegraphai/smart-scraper'.

Citations: [scrapegraphai/graphs/__init__.py:32-64](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/__init__.py#L32-L64) · [scrapegraphai/graphs/smart_scraper_graph.py:58-65](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/smart_scraper_graph.py#L58-L65) · [scrapegraphai/graphs/smart_scraper_graph.py:286-297](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/smart_scraper_graph.py#L286-L297) · [scrapegraphai/graphs/abstract_graph.py:56-108](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/abstract_graph.py#L56-L108) · [scrapegraphai/graphs/base_graph.py:344-376](https://github.com/ScrapeGraphAI/Scrapegraph-ai/blob/194055e203afce41ed4e70365dbc416bad756115/scrapegraphai/graphs/base_graph.py#L344-L376)
