ScrapeGraphAI/Scrapegraph-ai
Python library that turns a prompt and a URL into JSON by running LangChain LLMs over fetched, chunked pages in node graphs.
Overview
ScrapeGraphAI is a Python library that writes no scraping logic by hand: you give it a natural-language prompt, a source (URL, local file or raw text) and an LLM config, and it returns JSON. Each task is a “graph”, a small directed pipeline of nodes such as fetch, parse and generate-answer. SmartScraperGraph handles one page. SearchGraph searches the web first. SmartScraperMultiGraph fans out over URLs. There are also variants for CSV, JSON and XML, depth-limited crawling, screenshots, speech and script generation.
The engineering weight is in prompting and orchestration, not fetching. Pages are loaded with Playwright (plus a stealth patch), converted to text or Markdown, cut into token-sized chunks, and sent to any LangChain chat model. With several chunks, each chunk is answered in parallel and a final “merge” prompt combines the partial answers. Without a schema, output is free-form JSON. With a Pydantic schema, the model gets format instructions and the reply is parsed against them.
It is a good fit for quick prompt-driven extraction from a handful of pages. It is not a crawler framework. There is no persistent queue, no robots.txt enforcement in the built-in graphs, and no retry or proxy strategy beyond what you configure on the browser loader.
Architecture
flowchart LR
U["SmartScraperGraph(prompt, source, config)"] --> AG["AbstractGraph: _create_llm"]
AG --> BG["BaseGraph.execute"]
BG --> FN["FetchNode"]
FN --> CL["ChromiumLoader (Playwright)"]
FN --> ALT["BrowserBase / ScrapeDo / Plasmate"]
FN --> PN["ParseNode: html2text + semchunk"]
PN --> GA["GenerateAnswerNode"]
GA --> LLM["LangChain chat model"]
GA --> MG["merge prompt"]
BG --> TL["telemetry (on by default)"]
SG["SearchGraph"] --> SI["SearchInternetNode"]
SI --> GI["GraphIteratorNode"]
GI --> U
| Component | Path | Role |
|---|---|---|
| Graph base | scrapegraphai/graphs/abstract_graph.py |
Builds the LLM from config, sets token budget, pushes common params to nodes |
| Executor | scrapegraphai/graphs/base_graph.py |
Walks nodes along edges, tracks token cost per node, sends telemetry |
| Graphs | scrapegraphai/graphs/*.py |
About 25 ready pipelines (SmartScraperGraph, SearchGraph, DepthSearchGraph, …) |
| Nodes | scrapegraphai/nodes/ |
Fetch, parse, generate-answer, conditional, search, iterator, merge, code generation |
| Loaders | scrapegraphai/docloaders/ |
ChromiumLoader (Playwright or Selenium), BrowserBase, ScrapeDo, Plasmate |
| Prompts | scrapegraphai/prompts/ |
Templates for single-chunk, per-chunk and merge calls (HTML and Markdown variants) |
| Utils | scrapegraphai/utils/ |
convert_to_md, cleanup_html, split_text_into_chunks, output parsers, proxy broker, web search |
| Models | scrapegraphai/models/, helpers/models_tokens.py |
Wrappers for non-LangChain providers; context-window table per model |
| Telemetry | scrapegraphai/telemetry/telemetry.py |
Posts run data to the maintainers’ tracing endpoint |
How a request flows
Take SmartScraperGraph("List the products", "https://example.com", {"llm": {"model": "openai/gpt-4o-mini"}}).run():
- Build the LLM.
AbstractGraph.__init__calls_create_llm. It splitsprovider/model, or guesses the provider from the token table, checks it against a fixed allow-list, and setsmodel_tokenfrommodels_tokens(falling back to 8192 with a warning). It then calls LangChain’sinit_chat_modelor a custom wrapper for DeepSeek, xAI, Nvidia and others (abstract_graph.py). - Pick a graph shape.
_create_graphchooses one of eight node layouts from thehtml_mode,reasoningandreattemptflags. The default isFetchNode -> ParseNode -> GenerateAnswerNode(smart_scraper_graph.py). - Execute.
BaseGraph._execute_standardstarts at the entry node and runs each node inside a token-counting callback. It followsedges, or the node name aConditionalNodereturns (base_graph.py). - Fetch.
FetchNode.handle_web_sourceuses BrowserBase, ScrapeDo or Plasmate if configured. Otherwise it usesChromiumLoader. For OpenAI and Azure models (or withforce), the HTML is converted to Markdown withhtml2text(fetch_node.py). - Load the page.
ChromiumLoader.ascrape_playwrightlaunches Chromium (or Firefox) with an optional proxy, appliesMalenia.apply_stealthfromundetected-playwrightto the context, goes to the URL withdomcontentloaded, waits forload_stateand returnspage.content(), retrying up toretry_limit(default 1) (chromium.py). - Parse and chunk.
ParseNoderuns LangChain’sHtml2TextTransformer(links kept) and splits the text withsemchunktomodel_token - 250tokens. It also warns when none of the prompt’s or schema’s terms appear in the text (parse_node.py). - Answer.
GenerateAnswerNodebuilds format instructions from the schema, or asks for{"content": ...}. One chunk means one call. Several chunks become aRunnableParallelof per-chunk chains, then one merge call over all partial answers (generate_answer_node.py).run()returnsstate["answer"].
Key components
Chunking and merging
split_text_into_chunks uses semchunk with a token counter and a further 10% safety margin. A word-based splitter is kept as a fallback (split_text_into_chunks.py). The map-then-merge pattern scales to long pages. But each chunk’s answer is produced without seeing the others, and the merge step is one more unconstrained LLM call that can drop or invent rows.
Structured output
Schema support is prompt-level. get_pydantic_output_parser returns a LangChain JsonOutputParser for Pydantic v2 models and rejects v1 (output_parser.py). The schema becomes format instructions in the prompt, and the reply is parsed as JSON. Provider-side constrained decoding is used only for Ollama, where the node sets llm_model.format to the schema’s JSON Schema (generate_answer_node.py). Bedrock gets no format instructions at all.
Multi-page graphs
SearchGraph chains SearchInternetNode (DuckDuckGo by default, or Bing, SearXNG or Serper), GraphIteratorNode and MergeAnswersNode (search_graph.py). The iterator creates one SmartScraperGraph per URL and runs each graph.run in a thread under an asyncio.Semaphore (default batch size 16) (graph_iterator_node.py). DepthSearchGraph uses FetchNodeLevelK, which repeats link extraction for depth rounds over an in-memory document list (fetch_node_level_k.py).
robots.txt
RobotsNode fetches /robots.txt and asks the LLM, with a prompt, whether the path is allowed for an agent name looked up in helpers/robots.py. It raises unless force_scraping is set (robots_node.py). None of the shipped graphs include it, so in practice nothing checks robots.txt unless you build a custom graph.
Telemetry
After every successful run, BaseGraph calls log_graph_execution with the prompt, schema, parsed page content, LLM answer, model name and URL (base_graph.py). Telemetry is on by default. When all fields are present, which is any URL run with a schema, the payload is posted to an external tracing endpoint, capped at 1000 calls per session (telemetry.py, L177-L215). Set SCRAPEGRAPHAI_TELEMETRY_ENABLED=false if scraped content or prompts are sensitive.
Extending it
- Custom graphs. Subclass
AbstractGraph, implement_create_graphto return aBaseGraph(nodes, edges, entry_point), andrun. Nodes declare inputs with a small boolean expression language ("user_prompt & (relevant_chunks | parsed_doc | doc)") that picks keys from the shared state dict. - Custom nodes. Subclass
BaseNodeand implementexecute(state) -> state. - Any model. Pass
{"model_instance": your_langchain_chat_model, "model_tokens": N}to skip the provider allow-list (abstract_graph.py). - Fetch backends.
loader_kwargsreachChromiumLoader(backend="selenium",requires_js_support,load_state,proxy,retry_limit).browser_base,scrape_doandplasmateswitch to hosted fetchers. - Burr.
burr_kwargsruns the graph through the Burr state-machine framework for tracking and replay.
Running it
pip install scrapegraphai, thenplaywright installfor the browser. Python 3.12 or newer is required.- An LLM is mandatory: an API key for a hosted provider, or a local Ollama model (
"model": "ollama/llama3.2"). Setmodel_tokensfor models missing from the table, or long pages will be chunked against an 8192-token default. - There is no server, CLI or MCP endpoint in this repository. Everything runs in-process. The
docker-compose.ymlonly bundles an Ollama container.
Strengths and caveats
- Strength: very little code per task. A prompt and an optional Pydantic model replace selectors, and the graph library covers common shapes (single page, search, multi-URL, files).
- Strength: provider-agnostic. Anything LangChain can instantiate works, including local models through Ollama.
- Strength: clear pipeline model. Nodes and a shared state dict make it easy to add a reasoning step, a retry branch or a custom node.
- Caveat: an LLM call on every page. Cost and latency scale with page size and chunk count. Nothing generates reusable selectors for the default scrape path.
- Caveat: telemetry ships content off-box by default. Prompts, page text and answers go to the maintainers’ endpoint unless you disable it.
- Caveat: thin fetching. One retry attempt by default, a basic stealth patch, no block detection and no robots.txt in the shipped graphs. The
use_soupHTTP path inFetchNodealso references variables it never sets, so it fails if enabled. - Caveat: dead cloud shortcut.
SmartScraperGraphcompares the model object to the string"scrapegraphai/smart-scraper"to forward to the hosted API. That branch can never run, because_create_llmrejects thescrapegraphaiprovider first.
Sources: code at 194055e, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (12 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 194055e. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredPlain HTTP vs. headless browser. FetchNode has two paths controlled by the use_soup config flag (default False). When use_soup=True, it uses requests.get (plain HTTP, no JS). When use_soup=False, it delegates to ChromiumLoader, which spins up a real headless Chromium via Playwright (or Selenium/undetected-chromedriver with the backend parameter).
JS rendering. The ChromiumLoader has two scraping modes: ascrape_playwright waits for "domcontentloaded", while ascrape_with_js_support waits until "networkidle", ensuring JavaScript-rendered content is fully loaded. The FetchNode also passes requires_js_support config to control this. There is also an ascrape_playwright_scroll variant that uses mouse-wheel events and configurable sleep intervals for lazy-loaded/infinite-scroll pages.
Waiting strategy. The load_state parameter (default "domcontentloaded") controls how long Playwright waits before capturing page content. The node also supports retry_limit (attemps) and timeout (blocking-ops seconds).
Supported content types. Beyond HTML, FetchNode.handle_file loads PDFs (via PyPDFLoader), CSVs (pandas), JSON, XML, and Markdown files from local paths. It also handles entire directories of these types, and local HTML strings via handle_local_source. Search-engine results from SearchInternetNode return plain lists of URLs.
Third-party fetching services. The node supports browser_base (BrowserBase API), scrape_do (ScrapeDo proxy API), and plasmate (PlasmateLoader) as alternative fetching backends.
How is content extracted or converted?
answeredHTML→Markdown conversion. ParseNode.execute uses LangChain's Html2TextTransformer to convert HTML to plain text with links preserved. Separately, FetchNode conditionally calls convert_to_md() which uses the html2text library to produce Markdown from raw HTML, enabling either full-Markdown or text-only downstream processing. The cut=True flag controls whether the cleanup step is applied.
Boilerplate removal / HTML cleanup. The cleanup_html() function (BeautifulSoup-based) extracts the <title>, removes <style> tags, minifies body HTML via the minify library, extracts JSON from <script> tags, and collects link/image URLs. A separate reduce_html() function applies increasing levels of reduction: level 0 = regex minification, level 1 = remove comments and non-essential attributes, level 2 = truncate text content to 20 chars per tag. Default filters in helpers/default_filters.py provide common image-extension lists.
Content-chunking with schema awareness. After HTML→text conversion, ParseNode calls split_text_into_chunks() which uses the semchunk library (or word-level fallback) to split content into token-budgeted chunks sized to the LLM's model_token limit. The chunk_size is reduced by 250 (HTML mode) or 500 (other modes) to leave room for prompt instructions. The parser also collects URL fields that contain [link URLs, image URLs] for schema-aware follow-up.
Schema-based extraction. Extraction shaping is done downstream in GenerateAnswerNode. If the user provides a Pydantic schema, the LLM receives format instructions via get_pydantic_output_parser, which forces structured JSON output matching the schema. Without a schema, a TolerantJsonOutputParser is used with a default JSON format instruction.
Content-evidence warning. ParseNode has a _warn_if_content_lacks_requested_fields method that deterministically checks whether the user's prompt terms or schema field names appear in the parsed text — flagging empty/error-page/dropped-content conditions before they reach the LLM.
How are LLMs used, if at all?
answeredExtensive LLM usage across every pipeline stage. The library uses LLMs not just for final answer generation but also for search-query formulation (SearchInternetNode), HTML description (DescriptionNode), code generation (GenerateCodeNode), prompt refinement (PromptRefinerNode), URL relevance ranking (SearchLinkNode), and decision routing (ConditionalNode). Every GenerateAnswerNode chains a LangChain PromptTemplate to an LLM.
Prompting. All prompts come from scrapegraphai/prompts/ — distinct templates exist for single-chunk (TEMPLATE_NO_CHUNKS/TEMPLATE_NO_CHUNKS_MD), multi-chunk (TEMPLATE_CHUNKS/TEMPLATE_CHUNKS_MD), and merge (TEMPLATE_MERGE/TEMPLATE_MERGE_MD) workflows. Additional info can be prepended. When is_md_scraper or not script_creator is true, the Markdown variants (which ask the LLM to reason over Markdown-syntax content) are used.
Chunking large pages. GenerateAnswerNode.execute checks len(doc): a single chunk goes directly to the LLM; multiple chunks are processed concurrently via RunnableParallel (one chain per chunk), then a final merge prompt sends all partial results to the LLM for synthesis. Chunk size is driven by the model's max-token cap from helpers/models_tokens.py.
Structured output / JSON schema. When a Pydantic schema is provided: for ChatOpenAI models it uses get_pydantic_output_parser (which binds response_format or function-calling). For ChatOllama, it sets llm_model.format to the schema's JSON schema. For ChatBedrock, no format instructions are passed. Without a schema, TolerantJsonOutputParser wraps the response in {"content": ...}.
Supported providers. The _create_llm method in AbstractGraph supports 20+ providers: OpenAI, Azure, Google Gemini/Vertex, Ollama, Anthropic, Bedrock, MistralAI, Groq, Hugging Face, DeepSeek, TogetherAI, Fireworks, Ernie, plus custom models (CLoD, DeepSeek, MiniMax, Nvidia, OneApi, XAI). Provider is auto-detected from model name or specified via provider/model format.
Cost controls. AbstractGraph supports InMemoryRateLimiter (requests per second) and max_retries for rate limiting. Costs are tracked per-model in utils/model_costs.py (input cost per 1K tokens). The BaseGraph execution info reports total tokens, prompt/completion tokens, and cumulative USD cost per run. The model_tokens config parameter lets users override the default window (falls back to 8192 with a warning).
How are anti-bot measures, proxies and fingerprinting handled?
answeredStealth patches. The ChromiumLoader uses undetected-playwright's Malenia.apply_stealth() on every browser context (chromium.py:392). This patches Playwright's browser fingerprint to avoid bot detection. For Selenium, it uses undetected_chromedriver (uc) which patches the ChromeDriver to bypass anti-bot systems.
User-agent rotation. research_web.py maintains a hardcoded list of 5 real-looking user-agent strings (Chrome, Safari, Firefox, mobile variants) and selects one randomly via get_random_user_agent() for search-engine HTTP requests. This is separate from the browser backend.
Proxy rotation. The utils/proxy_rotation.py module provides full proxy brokering: search_proxy_servers() uses the free-proxy library (FreeProxy) to discover proxies matching criteria (anonymity, country set, HTTPS support, timeout). It tests each candidate by making a request and returns verified working proxies. The parse_or_search_proxy() function routes to either a known proxy server or the broker. Proxies are passed to Playwright's browser launch() method as the proxy parameter.
Third-party proxy services. FetchNode supports proxy-gated fetching via scrape_do (ScrapeDo API with geoCode/superProxy options) and browser_base (BrowserBase cloud browser API). PlasmateLoader is another alternative backend with its own header/proxy settings.
CAPTCHA handling. There is no explicit CAPTCHA solving module in the codebase. The combination of undetected-playwright stealth patches and undetected-chromedriver is the primary defense against CAPTCHA triggers. storage_state can persist authenticated browser sessions to avoid re-auth.
Rate limiting. research_web.py has a @rate_limited decorator (default: 10 calls per 60 seconds) applied to search_on_web for polite search-engine requests. LLM calls can be rate-limited via LangChain's InMemoryRateLimiter.
HTTP error detection. Both the use_soup path (response.status_code check) and ChromiumLoader (_warn_on_error_status) detect 4xx/5xx responses and log warnings, preventing error-page content from silently reaching the LLM.
How is crawling at scale implemented?
answeredLink-level crawling (depth-limited). The DepthSearchGraph pipeline uses FetchNodeLevelK, which implements recursive BFS-style crawling: it fetches a URL, extracts all <a href=...> links via BeautifulSoup, resolves relative URLs, deduplicates against already-seen URLs, then fetches the new links. This repeats for depth iterations. The only_inside_links flag restricts to same-domain links. The FetchNodeLevelK.obtain_content() method manages the frontier as a document list where each entry is {"source": url, "document": ...}.
No robots.txt enforcement. Despite a helpers/robots.py file, it only maps model names to bot user-agent strings (e.g., "gpt-4o"→"GPTBot"). There is no actual robots.txt fetching, parsing, or politeness-delay implementation in the code. A robots_node.py module exists but bundles an LLM prompt rather than implementing robots.txt parsing.
Concurrency via asyncio semaphore. The GraphIteratorNode runs multiple graph instances concurrently using asyncio.Semaphore(batchsize) (default 16). For each URL in the input list, it spawns a graph.run() call in a thread via asyncio.to_thread, and limits parallel execution with the semaphore. Tqdm progress bars track completion.
URL deduplication. In FetchNodeLevelK.obtain_content(), newly extracted links are checked against both the documents and new_documents lists before being added to the frontier: if not any(d.get("source") == link for d in documents) and not any(d.get("source") == link for d in new_documents). There is no shared visited set across graph runs.
No distributed workers. The codebase has no message-queue, Redis, or distributed-worker infrastructure. All concurrency is in-process via asyncio. The batch_api.py utility hints at batched LLM API calls but does not distribute scraping work.
Max results limits. Graphs accept max_results (default 3 in SearchGraph) and max_results in SearchInternetNode to cap the number of fetched pages. The AbstractGraph config can set max_results for overall graph iteration counts in graph memory.
What is the developer interface?
answeredLibrary API (primary). The main developer interface is importing graph classes from scrapegraphai.graphs and running them: SmartScraperGraph(prompt, source, config).run(). The 25+ graph variants handle different use cases — single-page (SmartScraperGraph), multi-page (SmartScraperMultiGraph), search-engine (SearchGraph), depth-recursive (DepthSearchGraph), CSV/JSON/XML (CSVScraperGraph, JSONScraperGraph, XMLScraperGraph), screenshot (ScreenshotScraperGraph), image-to-text (OmniScraperGraph), speech output (SpeechGraph), and code-generation (ScriptCreatorGraph). Config is a plain dict with llm, headless, verbose, etc.
Graph builder. The builders/graph_builder.py provides GraphBuilder for constructing custom pipelines from YAML configs without writing Python code.
Burr integration. The integrations/burr_bridge.py allows executing graphs via the Burr orchestration framework when burr_kwargs is passed in config, providing state-persistence and replay capabilities.
CLI. There is no dedicated CLI entry point in pyproject.toml or setup files. The library is used programmatically. However, the integrations include a scrapegraph_py_compat module that proxies SmartScraperGraph to the cloud ScrapeGraphAI service when llm_model == "scrapegraphai/smart-scraper".
Output formats. Results are Python dicts (JSON). The utils/data_export.py supports CSV/JSON export. utils/save_audio_from_bytes.py handles audio output for speech graphs. utils/save_code_to_file.py saves generated scraping scripts.
Language bindings. Python only. The pyproject.toml requires Python ≥3.12. No REST service, MCP server, or UI is included in the open-source repository. The scrapegraph-py SDK dependency (≥2.0.0) connects to https://scrapegraphai.com cloud API when configured, but the open-source code runs locally via LangChain and Playwright.
Error reporting. The execution info dict returned by BaseGraph includes per-node token counts, execution time, and cumulative USD cost, accessible via graph.get_execution_info().