LLMs Technical Reviews
Home / AI web scraping / Scrapegraph-ai

ScrapeGraphAI/Scrapegraph-ai

Python library that turns a prompt and a URL into JSON by running LangChain LLMs over fetched, chunked pages in node graphs.

GitHub ↗★ 32kPythonMITcommit 194055e · 2026-10-06homepage ↗

Overview

ScrapeGraphAI is a Python library that writes no scraping logic by hand: you give it a natural-language prompt, a source (URL, local file or raw text) and an LLM config, and it returns JSON. Each task is a “graph”, a small directed pipeline of nodes such as fetch, parse and generate-answer. SmartScraperGraph handles one page. SearchGraph searches the web first. SmartScraperMultiGraph fans out over URLs. There are also variants for CSV, JSON and XML, depth-limited crawling, screenshots, speech and script generation.

The engineering weight is in prompting and orchestration, not fetching. Pages are loaded with Playwright (plus a stealth patch), converted to text or Markdown, cut into token-sized chunks, and sent to any LangChain chat model. With several chunks, each chunk is answered in parallel and a final “merge” prompt combines the partial answers. Without a schema, output is free-form JSON. With a Pydantic schema, the model gets format instructions and the reply is parsed against them.

It is a good fit for quick prompt-driven extraction from a handful of pages. It is not a crawler framework. There is no persistent queue, no robots.txt enforcement in the built-in graphs, and no retry or proxy strategy beyond what you configure on the browser loader.

Architecture

flowchart LR
  U["SmartScraperGraph(prompt, source, config)"] --> AG["AbstractGraph: _create_llm"]
  AG --> BG["BaseGraph.execute"]
  BG --> FN["FetchNode"]
  FN --> CL["ChromiumLoader (Playwright)"]
  FN --> ALT["BrowserBase / ScrapeDo / Plasmate"]
  FN --> PN["ParseNode: html2text + semchunk"]
  PN --> GA["GenerateAnswerNode"]
  GA --> LLM["LangChain chat model"]
  GA --> MG["merge prompt"]
  BG --> TL["telemetry (on by default)"]
  SG["SearchGraph"] --> SI["SearchInternetNode"]
  SI --> GI["GraphIteratorNode"]
  GI --> U
Component Path Role
Graph base scrapegraphai/graphs/abstract_graph.py Builds the LLM from config, sets token budget, pushes common params to nodes
Executor scrapegraphai/graphs/base_graph.py Walks nodes along edges, tracks token cost per node, sends telemetry
Graphs scrapegraphai/graphs/*.py About 25 ready pipelines (SmartScraperGraph, SearchGraph, DepthSearchGraph, …)
Nodes scrapegraphai/nodes/ Fetch, parse, generate-answer, conditional, search, iterator, merge, code generation
Loaders scrapegraphai/docloaders/ ChromiumLoader (Playwright or Selenium), BrowserBase, ScrapeDo, Plasmate
Prompts scrapegraphai/prompts/ Templates for single-chunk, per-chunk and merge calls (HTML and Markdown variants)
Utils scrapegraphai/utils/ convert_to_md, cleanup_html, split_text_into_chunks, output parsers, proxy broker, web search
Models scrapegraphai/models/, helpers/models_tokens.py Wrappers for non-LangChain providers; context-window table per model
Telemetry scrapegraphai/telemetry/telemetry.py Posts run data to the maintainers’ tracing endpoint

How a request flows

Take SmartScraperGraph("List the products", "https://example.com", {"llm": {"model": "openai/gpt-4o-mini"}}).run():

  1. Build the LLM. AbstractGraph.__init__ calls _create_llm. It splits provider/model, or guesses the provider from the token table, checks it against a fixed allow-list, and sets model_token from models_tokens (falling back to 8192 with a warning). It then calls LangChain’s init_chat_model or a custom wrapper for DeepSeek, xAI, Nvidia and others (abstract_graph.py).
  2. Pick a graph shape. _create_graph chooses one of eight node layouts from the html_mode, reasoning and reattempt flags. The default is FetchNode -> ParseNode -> GenerateAnswerNode (smart_scraper_graph.py).
  3. Execute. BaseGraph._execute_standard starts at the entry node and runs each node inside a token-counting callback. It follows edges, or the node name a ConditionalNode returns (base_graph.py).
  4. Fetch. FetchNode.handle_web_source uses BrowserBase, ScrapeDo or Plasmate if configured. Otherwise it uses ChromiumLoader. For OpenAI and Azure models (or with force), the HTML is converted to Markdown with html2text (fetch_node.py).
  5. Load the page. ChromiumLoader.ascrape_playwright launches Chromium (or Firefox) with an optional proxy, applies Malenia.apply_stealth from undetected-playwright to the context, goes to the URL with domcontentloaded, waits for load_state and returns page.content(), retrying up to retry_limit (default 1) (chromium.py).
  6. Parse and chunk. ParseNode runs LangChain’s Html2TextTransformer (links kept) and splits the text with semchunk to model_token - 250 tokens. It also warns when none of the prompt’s or schema’s terms appear in the text (parse_node.py).
  7. Answer. GenerateAnswerNode builds format instructions from the schema, or asks for {"content": ...}. One chunk means one call. Several chunks become a RunnableParallel of per-chunk chains, then one merge call over all partial answers (generate_answer_node.py). run() returns state["answer"].

Key components

Chunking and merging

split_text_into_chunks uses semchunk with a token counter and a further 10% safety margin. A word-based splitter is kept as a fallback (split_text_into_chunks.py). The map-then-merge pattern scales to long pages. But each chunk’s answer is produced without seeing the others, and the merge step is one more unconstrained LLM call that can drop or invent rows.

Structured output

Schema support is prompt-level. get_pydantic_output_parser returns a LangChain JsonOutputParser for Pydantic v2 models and rejects v1 (output_parser.py). The schema becomes format instructions in the prompt, and the reply is parsed as JSON. Provider-side constrained decoding is used only for Ollama, where the node sets llm_model.format to the schema’s JSON Schema (generate_answer_node.py). Bedrock gets no format instructions at all.

Multi-page graphs

SearchGraph chains SearchInternetNode (DuckDuckGo by default, or Bing, SearXNG or Serper), GraphIteratorNode and MergeAnswersNode (search_graph.py). The iterator creates one SmartScraperGraph per URL and runs each graph.run in a thread under an asyncio.Semaphore (default batch size 16) (graph_iterator_node.py). DepthSearchGraph uses FetchNodeLevelK, which repeats link extraction for depth rounds over an in-memory document list (fetch_node_level_k.py).

robots.txt

RobotsNode fetches /robots.txt and asks the LLM, with a prompt, whether the path is allowed for an agent name looked up in helpers/robots.py. It raises unless force_scraping is set (robots_node.py). None of the shipped graphs include it, so in practice nothing checks robots.txt unless you build a custom graph.

Telemetry

After every successful run, BaseGraph calls log_graph_execution with the prompt, schema, parsed page content, LLM answer, model name and URL (base_graph.py). Telemetry is on by default. When all fields are present, which is any URL run with a schema, the payload is posted to an external tracing endpoint, capped at 1000 calls per session (telemetry.py, L177-L215). Set SCRAPEGRAPHAI_TELEMETRY_ENABLED=false if scraped content or prompts are sensitive.

Extending it

  • Custom graphs. Subclass AbstractGraph, implement _create_graph to return a BaseGraph(nodes, edges, entry_point), and run. Nodes declare inputs with a small boolean expression language ("user_prompt & (relevant_chunks | parsed_doc | doc)") that picks keys from the shared state dict.
  • Custom nodes. Subclass BaseNode and implement execute(state) -> state.
  • Any model. Pass {"model_instance": your_langchain_chat_model, "model_tokens": N} to skip the provider allow-list (abstract_graph.py).
  • Fetch backends. loader_kwargs reach ChromiumLoader (backend="selenium", requires_js_support, load_state, proxy, retry_limit). browser_base, scrape_do and plasmate switch to hosted fetchers.
  • Burr. burr_kwargs runs the graph through the Burr state-machine framework for tracking and replay.

Running it

  • pip install scrapegraphai, then playwright install for the browser. Python 3.12 or newer is required.
  • An LLM is mandatory: an API key for a hosted provider, or a local Ollama model ("model": "ollama/llama3.2"). Set model_tokens for models missing from the table, or long pages will be chunked against an 8192-token default.
  • There is no server, CLI or MCP endpoint in this repository. Everything runs in-process. The docker-compose.yml only bundles an Ollama container.

Strengths and caveats

  • Strength: very little code per task. A prompt and an optional Pydantic model replace selectors, and the graph library covers common shapes (single page, search, multi-URL, files).
  • Strength: provider-agnostic. Anything LangChain can instantiate works, including local models through Ollama.
  • Strength: clear pipeline model. Nodes and a shared state dict make it easy to add a reasoning step, a retry branch or a custom node.
  • Caveat: an LLM call on every page. Cost and latency scale with page size and chunk count. Nothing generates reusable selectors for the default scrape path.
  • Caveat: telemetry ships content off-box by default. Prompts, page text and answers go to the maintainers’ endpoint unless you disable it.
  • Caveat: thin fetching. One retry attempt by default, a basic stealth patch, no block detection and no robots.txt in the shipped graphs. The use_soup HTTP path in FetchNode also references variables it never sets, so it fails if enabled.
  • Caveat: dead cloud shortcut. SmartScraperGraph compares the model object to the string "scrapegraphai/smart-scraper" to forward to the hosted API. That branch can never run, because _create_llm rejects the scrapegraphai provider first.

Sources: code at 194055e, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (12 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 194055e. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Plain HTTP vs. headless browser. FetchNode has two paths controlled by the use_soup config flag (default False). When use_soup=True, it uses requests.get (plain HTTP, no JS). When use_soup=False, it delegates to ChromiumLoader, which spins up a real headless Chromium via Playwright (or Selenium/undetected-chromedriver with the backend parameter).

JS rendering. The ChromiumLoader has two scraping modes: ascrape_playwright waits for "domcontentloaded", while ascrape_with_js_support waits until "networkidle", ensuring JavaScript-rendered content is fully loaded. The FetchNode also passes requires_js_support config to control this. There is also an ascrape_playwright_scroll variant that uses mouse-wheel events and configurable sleep intervals for lazy-loaded/infinite-scroll pages.

Waiting strategy. The load_state parameter (default "domcontentloaded") controls how long Playwright waits before capturing page content. The node also supports retry_limit (attemps) and timeout (blocking-ops seconds).

Supported content types. Beyond HTML, FetchNode.handle_file loads PDFs (via PyPDFLoader), CSVs (pandas), JSON, XML, and Markdown files from local paths. It also handles entire directories of these types, and local HTML strings via handle_local_source. Search-engine results from SearchInternetNode return plain lists of URLs.

Third-party fetching services. The node supports browser_base (BrowserBase API), scrape_do (ScrapeDo proxy API), and plasmate (PlasmateLoader) as alternative fetching backends.

How is content extracted or converted?

answered

HTML→Markdown conversion. ParseNode.execute uses LangChain's Html2TextTransformer to convert HTML to plain text with links preserved. Separately, FetchNode conditionally calls convert_to_md() which uses the html2text library to produce Markdown from raw HTML, enabling either full-Markdown or text-only downstream processing. The cut=True flag controls whether the cleanup step is applied.

Boilerplate removal / HTML cleanup. The cleanup_html() function (BeautifulSoup-based) extracts the <title>, removes <style> tags, minifies body HTML via the minify library, extracts JSON from <script> tags, and collects link/image URLs. A separate reduce_html() function applies increasing levels of reduction: level 0 = regex minification, level 1 = remove comments and non-essential attributes, level 2 = truncate text content to 20 chars per tag. Default filters in helpers/default_filters.py provide common image-extension lists.

Content-chunking with schema awareness. After HTML→text conversion, ParseNode calls split_text_into_chunks() which uses the semchunk library (or word-level fallback) to split content into token-budgeted chunks sized to the LLM's model_token limit. The chunk_size is reduced by 250 (HTML mode) or 500 (other modes) to leave room for prompt instructions. The parser also collects URL fields that contain [link URLs, image URLs] for schema-aware follow-up.

Schema-based extraction. Extraction shaping is done downstream in GenerateAnswerNode. If the user provides a Pydantic schema, the LLM receives format instructions via get_pydantic_output_parser, which forces structured JSON output matching the schema. Without a schema, a TolerantJsonOutputParser is used with a default JSON format instruction.

Content-evidence warning. ParseNode has a _warn_if_content_lacks_requested_fields method that deterministically checks whether the user's prompt terms or schema field names appear in the parsed text — flagging empty/error-page/dropped-content conditions before they reach the LLM.

How are LLMs used, if at all?

answered

Extensive LLM usage across every pipeline stage. The library uses LLMs not just for final answer generation but also for search-query formulation (SearchInternetNode), HTML description (DescriptionNode), code generation (GenerateCodeNode), prompt refinement (PromptRefinerNode), URL relevance ranking (SearchLinkNode), and decision routing (ConditionalNode). Every GenerateAnswerNode chains a LangChain PromptTemplate to an LLM.

Prompting. All prompts come from scrapegraphai/prompts/ — distinct templates exist for single-chunk (TEMPLATE_NO_CHUNKS/TEMPLATE_NO_CHUNKS_MD), multi-chunk (TEMPLATE_CHUNKS/TEMPLATE_CHUNKS_MD), and merge (TEMPLATE_MERGE/TEMPLATE_MERGE_MD) workflows. Additional info can be prepended. When is_md_scraper or not script_creator is true, the Markdown variants (which ask the LLM to reason over Markdown-syntax content) are used.

Chunking large pages. GenerateAnswerNode.execute checks len(doc): a single chunk goes directly to the LLM; multiple chunks are processed concurrently via RunnableParallel (one chain per chunk), then a final merge prompt sends all partial results to the LLM for synthesis. Chunk size is driven by the model's max-token cap from helpers/models_tokens.py.

Structured output / JSON schema. When a Pydantic schema is provided: for ChatOpenAI models it uses get_pydantic_output_parser (which binds response_format or function-calling). For ChatOllama, it sets llm_model.format to the schema's JSON schema. For ChatBedrock, no format instructions are passed. Without a schema, TolerantJsonOutputParser wraps the response in {"content": ...}.

Supported providers. The _create_llm method in AbstractGraph supports 20+ providers: OpenAI, Azure, Google Gemini/Vertex, Ollama, Anthropic, Bedrock, MistralAI, Groq, Hugging Face, DeepSeek, TogetherAI, Fireworks, Ernie, plus custom models (CLoD, DeepSeek, MiniMax, Nvidia, OneApi, XAI). Provider is auto-detected from model name or specified via provider/model format.

Cost controls. AbstractGraph supports InMemoryRateLimiter (requests per second) and max_retries for rate limiting. Costs are tracked per-model in utils/model_costs.py (input cost per 1K tokens). The BaseGraph execution info reports total tokens, prompt/completion tokens, and cumulative USD cost per run. The model_tokens config parameter lets users override the default window (falls back to 8192 with a warning).

Editor's note. Correction: get_pydantic_output_parser returns a plain LangChain JsonOutputParser for every non-Bedrock model; it does not bind response_format or function calling. The schema only becomes format instructions in the prompt, and only Ollama gets provider-side constraint (llm_model.format set to the JSON Schema).

How are anti-bot measures, proxies and fingerprinting handled?

answered

Stealth patches. The ChromiumLoader uses undetected-playwright's Malenia.apply_stealth() on every browser context (chromium.py:392). This patches Playwright's browser fingerprint to avoid bot detection. For Selenium, it uses undetected_chromedriver (uc) which patches the ChromeDriver to bypass anti-bot systems.

User-agent rotation. research_web.py maintains a hardcoded list of 5 real-looking user-agent strings (Chrome, Safari, Firefox, mobile variants) and selects one randomly via get_random_user_agent() for search-engine HTTP requests. This is separate from the browser backend.

Proxy rotation. The utils/proxy_rotation.py module provides full proxy brokering: search_proxy_servers() uses the free-proxy library (FreeProxy) to discover proxies matching criteria (anonymity, country set, HTTPS support, timeout). It tests each candidate by making a request and returns verified working proxies. The parse_or_search_proxy() function routes to either a known proxy server or the broker. Proxies are passed to Playwright's browser launch() method as the proxy parameter.

Third-party proxy services. FetchNode supports proxy-gated fetching via scrape_do (ScrapeDo API with geoCode/superProxy options) and browser_base (BrowserBase cloud browser API). PlasmateLoader is another alternative backend with its own header/proxy settings.

CAPTCHA handling. There is no explicit CAPTCHA solving module in the codebase. The combination of undetected-playwright stealth patches and undetected-chromedriver is the primary defense against CAPTCHA triggers. storage_state can persist authenticated browser sessions to avoid re-auth.

Rate limiting. research_web.py has a @rate_limited decorator (default: 10 calls per 60 seconds) applied to search_on_web for polite search-engine requests. LLM calls can be rate-limited via LangChain's InMemoryRateLimiter.

HTTP error detection. Both the use_soup path (response.status_code check) and ChromiumLoader (_warn_on_error_status) detect 4xx/5xx responses and log warnings, preventing error-page content from silently reaching the LLM.

How is crawling at scale implemented?

answered

Link-level crawling (depth-limited). The DepthSearchGraph pipeline uses FetchNodeLevelK, which implements recursive BFS-style crawling: it fetches a URL, extracts all <a href=...> links via BeautifulSoup, resolves relative URLs, deduplicates against already-seen URLs, then fetches the new links. This repeats for depth iterations. The only_inside_links flag restricts to same-domain links. The FetchNodeLevelK.obtain_content() method manages the frontier as a document list where each entry is {"source": url, "document": ...}.

No robots.txt enforcement. Despite a helpers/robots.py file, it only maps model names to bot user-agent strings (e.g., "gpt-4o"→"GPTBot"). There is no actual robots.txt fetching, parsing, or politeness-delay implementation in the code. A robots_node.py module exists but bundles an LLM prompt rather than implementing robots.txt parsing.

Concurrency via asyncio semaphore. The GraphIteratorNode runs multiple graph instances concurrently using asyncio.Semaphore(batchsize) (default 16). For each URL in the input list, it spawns a graph.run() call in a thread via asyncio.to_thread, and limits parallel execution with the semaphore. Tqdm progress bars track completion.

URL deduplication. In FetchNodeLevelK.obtain_content(), newly extracted links are checked against both the documents and new_documents lists before being added to the frontier: if not any(d.get("source") == link for d in documents) and not any(d.get("source") == link for d in new_documents). There is no shared visited set across graph runs.

No distributed workers. The codebase has no message-queue, Redis, or distributed-worker infrastructure. All concurrency is in-process via asyncio. The batch_api.py utility hints at batched LLM API calls but does not distribute scraping work.

Max results limits. Graphs accept max_results (default 3 in SearchGraph) and max_results in SearchInternetNode to cap the number of fetched pages. The AbstractGraph config can set max_results for overall graph iteration counts in graph memory.

Editor's note. Correction: RobotsNode does fetch /robots.txt and asks the LLM whether the path is allowed (raising unless force_scraping), but no shipped graph includes it, so default runs never consult robots.txt.

What is the developer interface?

answered

Library API (primary). The main developer interface is importing graph classes from scrapegraphai.graphs and running them: SmartScraperGraph(prompt, source, config).run(). The 25+ graph variants handle different use cases — single-page (SmartScraperGraph), multi-page (SmartScraperMultiGraph), search-engine (SearchGraph), depth-recursive (DepthSearchGraph), CSV/JSON/XML (CSVScraperGraph, JSONScraperGraph, XMLScraperGraph), screenshot (ScreenshotScraperGraph), image-to-text (OmniScraperGraph), speech output (SpeechGraph), and code-generation (ScriptCreatorGraph). Config is a plain dict with llm, headless, verbose, etc.

Graph builder. The builders/graph_builder.py provides GraphBuilder for constructing custom pipelines from YAML configs without writing Python code.

Burr integration. The integrations/burr_bridge.py allows executing graphs via the Burr orchestration framework when burr_kwargs is passed in config, providing state-persistence and replay capabilities.

CLI. There is no dedicated CLI entry point in pyproject.toml or setup files. The library is used programmatically. However, the integrations include a scrapegraph_py_compat module that proxies SmartScraperGraph to the cloud ScrapeGraphAI service when llm_model == "scrapegraphai/smart-scraper".

Output formats. Results are Python dicts (JSON). The utils/data_export.py supports CSV/JSON export. utils/save_audio_from_bytes.py handles audio output for speech graphs. utils/save_code_to_file.py saves generated scraping scripts.

Language bindings. Python only. The pyproject.toml requires Python ≥3.12. No REST service, MCP server, or UI is included in the open-source repository. The scrapegraph-py SDK dependency (≥2.0.0) connects to https://scrapegraphai.com cloud API when configured, but the open-source code runs locally via LangChain and Playwright.

Error reporting. The execution info dict returned by BaseGraph includes per-node token counts, execution time, and cumulative USD cost, accessible via graph.get_execution_info().

Editor's note. Correction: the scrapegraph_py_compat cloud shortcut in SmartScraperGraph is unreachable: _create_llm rejects the 'scrapegraphai' provider first, and llm_model is a model object that never equals the string 'scrapegraphai/smart-scraper'.