LLMs Technical Reviews
Home / AI web scraping / CyberScraper-2077

itsOwen/CyberScraper-2077

Streamlit chat app that renders a page with Patchright, flattens it to text, and asks an LLM to answer or export it as JSON, CSV or SQL.

GitHub ↗★ 3.3kPythonMITcommit a260fa8 · 2026-09-27

Overview

CyberScraper 2077 is a Streamlit chat app for one-off scraping. You paste a URL into the chat, it renders the page with Patchright (an undetectable fork of Playwright), strips the HTML to plain text, and keeps that text in memory. Each later message is answered by an LLM over that text. If your message contains words like “csv”, “json”, “excel”, “sql” or “table”, the prompt asks the model for a bare JSON array, and the app turns it into a DataFrame, a CSV or Excel download, SQL INSERTs, an HTML table, or a Google Sheet.

It is a personal tool with a themed UI, not a library or a service. There is no CLI, no HTTP API and no crawler. The model never sees HTML structure, selectors or a schema, only the visible text, so extraction quality is whatever the LLM can infer from flattened text. In return, the setup is small: pick a model (OpenAI, Gemini, Ollama, or anything behind a LiteLLM proxy), paste a URL, ask for a table.

Two extras stand out: .onion URLs are fetched through a local Tor SOCKS proxy, and a -captcha suffix opens a visible browser and waits for you to solve the challenge by hand.

Architecture

flowchart LR
  U["User in Streamlit chat"] --> M["main.py UI"]
  M --> C["StreamlitWebScraperChat"]
  C --> X["WebExtractor.process_query"]
  X -->|"URL"| F["_fetch_url"]
  F -->|".onion"| T["TorScraper via SOCKS 9050"]
  F -->|"other"| P["PlaywrightScraper (Patchright)"]
  P --> PRE["_preprocess_content to text"]
  T --> PRE
  X -->|"question"| E["_extract_info"]
  E --> L["LLM: OpenAI, Gemini, Ollama, LiteLLM"]
  L --> FMT["_format_result"]
  FMT --> OUT["JSON, CSV, Excel, SQL, HTML, Sheets"]
Component Path Role
Streamlit app main.py Sidebar model picker, chat history in chat_history.json, result rendering and downloads
Sync bridge app/streamlit_web_scraper_chat.py Wraps WebExtractor and runs each message with asyncio.run
Orchestrator src/web_extractor.py URL detection, fetch, preprocessing, LLM calls, chunking, cache, output formatting
Browser scraper src/scrapers/playwright_scraper.py ScraperConfig, pooled Patchright browser, multi-page and CAPTCHA modes
Tor scraper src/scrapers/tor/ aiohttp over SOCKS5 for .onion hosts, Tor connectivity check
Model factory src/models.py, src/ollama_models.py LangChain chat models, LiteLLM proxy via ChatOpenAI, Ollama over its REST API
Prompt src/prompts.py One shared template with a data-export mode and a conversational mode
Sheets export src/utils/google_sheets_utils.py OAuth flow and upload of a DataFrame

How a request flows

Take two messages: https://example.com/products?page=1 1-3, then give me a csv of name and price.

  1. Bridge. StreamlitWebScraperChat.process_message runs WebExtractor.process_query inside asyncio.run, with a Streamlit placeholder as the progress callback (streamlit_web_scraper_chat.py). main.py builds the ScraperConfig from the sidebar, with delay_after_load=5 (main.py).
  2. Parse the message. process_query finds the first URL with a regex. The token after it is treated as a page spec if it contains only digits, commas and dashes, the next token as an optional URL pattern, and -captcha anywhere turns on CAPTCHA mode (web_extractor.py).
  3. Fetch. _fetch_url sends .onion hosts to TorScraper and everything else to PlaywrightScraper.fetch_content with proxy=None (web_extractor.py). With pages, detect_url_pattern finds the first numeric query parameter (page={page}) or path segment, and scrape_multiple_pages opens one fresh context per page under a semaphore of 5, each followed by a random 0.5 to 1.5 s pause (playwright_scraper.py, L776-L851).
  4. Flatten. The HTML of all pages is joined and passed through _preprocess_content: drop script, style, header, footer, nav and aside, drop comments and empty tags, then get_text() and collapse whitespace (web_extractor.py). A content hash resets the query cache when the page changes.
  5. Ask. The second message has no URL, so _extract_info runs. If the text fits in max_tokens - 1000 (128k for gpt-4.1-mini and gpt-4o-mini, 16,385 otherwise), it makes one call; otherwise it splits into 32,000-token chunks, calls the model per chunk and merges the JSON arrays (web_extractor.py). _call_model fills the shared prompt with the last 10 chat turns (each cut to 500 characters), the page text and the query (web_extractor.py).
  6. Format. _format_result keyword-matches the query again, pulls a JSON array out of the reply (direct parse, fenced block, or regex), and for “csv” returns a CSV string plus a DataFrame (web_extractor.py). main.py shows the table and download buttons.

Key components

Prompt and output modes

There is one prompt for every model. It sets a “netrunner” persona, says to return only a JSON array for export-style requests, to use "N/A" for missing fields and never to invent data, and to answer in plain text otherwise (prompts.py). Mode selection is plain substring matching on the user’s message, so “extract” or “table” anywhere in a question switches to JSON. The SQL formatter builds a CREATE TABLE with all columns as TEXT and escapes single quotes (web_extractor.py).

Model routing

WebExtractor.__init__ picks the backend from the model name: ollama: goes to OllamaModel, gemini- to ChatGoogleGenerativeAI, everything else to Models.get_model (web_extractor.py). The factory accepts four OpenAI chat names, text-* completion models, and litellm:<model>, which becomes a ChatOpenAI pointed at LITELLM_BASE_URL (models.py). Ollama is called directly at /api/generate with streaming, opening a fresh aiohttp session per call because each Streamlit message runs in a new event loop (ollama_models.py).

Browser and stealth

ScraperConfig exposes many toggles, but the effective stealth is Patchright itself plus real Chrome where available (channel="chrome" except on ARM64 Linux), no custom user agent or viewport, and one small permissions.query patch per page (playwright_scraper.py, L420-L482, L539-L567). The extra init script is defined but its calls are commented out. bypass_cloudflare() and simulate_human_behavior() exist but nothing calls them, and the bypass_cloudflare flag is never read (playwright_scraper.py). The “Use Current Browser” option launches your installed Chrome with a temp profile and --remote-debugging-port=9222, then connects over CDP (playwright_scraper.py).

CAPTCHA mode

With -captcha, the browser launches headed, navigates, and blocks on aioconsole.ainput() until you press Enter in the terminal that runs Streamlit, then scrapes the current page and any further pages in the same tab, and closes the browser (playwright_scraper.py). That works on a desktop, not in Docker.

Tor

TorManager opens one aiohttp session through socks5://127.0.0.1:9050 with headers chosen once from three Firefox user agents, checks check.torproject.org/api/ip before each fetch, and returns the raw response text (tor_manager.py). There is no JavaScript rendering on this path, and auto_renew_circuit in TorConfig is never used (tor_config.py).

Extending it

  • New model provider. Add a case in Models.get_model and a matching prefix in get_prompt_for_model, or put the provider behind a LiteLLM proxy and use litellm:<name> with no code change.
  • New fetcher. Subclass BaseScraper (async fetch_content and extract) and route to it in WebExtractor._fetch_url, the same way .onion URLs go to TorScraper.
  • Use it without Streamlit. WebExtractor(model_name=...) plus await process_query(url) and then await process_query("give me json of ...") works as a two-step library call, since it holds the page text as instance state.
  • Better extraction. The quickest win is to stop flattening to text: pass cleaned HTML or markdown into _preprocess_content, or add a schema to the prompt.

Running it

  • Local. Python 3.10+ (the code uses match), pip install -r requirements.txt, patchright install chromium (or chrome), set OPENAI_API_KEY or GOOGLE_API_KEY, then streamlit run main.py. Ollama is reached at OLLAMA_BASE_URL (default http://localhost:11434).
  • Docker. The image is Python 3.12 slim with Tor and Patchright Chromium; a generated run.sh starts Tor, waits for port 9050, and launches Streamlit on 8501 (Dockerfile). “Use Current Browser” and CAPTCHA mode do not work inside the container.

Strengths and caveats

  • Strength: low ceremony. Paste a URL, name the fields, download a CSV. Pagination by 1-5 and .onion support are built in.
  • Strength: model freedom. OpenAI, Gemini, any Ollama model, and anything behind a LiteLLM proxy share one prompt and one code path.
  • Strength: honest stealth defaults. It relies on Patchright and real Chrome instead of fragile JS spoofing, and leaves user agent and viewport alone.
  • Caveat: text-only context. Links, attributes, image URLs and table structure are lost before the LLM sees the page, and header, nav and footer content is always dropped.
  • Caveat: brittle large-page path. Chunked results are merged with a plain json.loads on each reply, so fenced or chatty replies are silently dropped, and a conversational question over a long page comes back as a JSON array or as [] (web_extractor.py).
  • Caveat: advertised anti-bot features are idle. Cloudflare retry, human simulation, the init script and Tor circuit renewal are not wired in, and the app always passes proxy=None, so there is no proxy support for normal sites.
  • Caveat: single-user app. Chat history is a local chat_history.json, each message runs in its own event loop, and CAPTCHA mode waits on the server’s terminal.

Sources: code at a260fa8, deepwiki-open wiki (10 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit a260fa8. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Pages are fetched primarily via Patchright (undetected Playwright fork) for full headless browser automation with JavaScript rendering. The PlaywrightScraper class in src/scrapers/playwright_scraper.py orchestrates this: it launches a Chromium browser (or real Chrome on x86 platforms), creates browser contexts, navigates via page.goto() with a configurable wait strategy (default domcontentloaded in ScraperConfig, line 50), and extracts raw HTML via page.content(). A 'Current Browser' mode at line 274 launches the user's real Chrome with remote debugging on port 9222 and connects over CDP, using the user's own profile for maximum undetectability. For .onion URLs, TorScraper in src/scrapers/tor/tor_scraper.py routes via SOCKS5 at 127.0.0.1:9050 using aiohttp_socks.ProxyConnector (tor_manager.py line 53-54). Multi-page fetching supports URL pattern detection in query parameters or path segments (lines 776-802) and concurrent scraping via asyncio.gather with a configurable semaphore cap (max_concurrent_pages, line 650). A shared aiohttp ClientSession singleton in src/http_client.py provides plain-HTTP connection-pooled fetching for non-browser use. There is no PDF rendering, no image content extraction — the scraper extracts only HTML content from the browser.

Editor's note. Correction: src/http_client.py is not imported anywhere, so there is no plain-HTTP fetch path; non-onion pages always go through Patchright and onion pages through TorManager's own aiohttp session.

How is content extracted or converted?

answered

Extraction is a two-stage pipeline. First, raw HTML is preprocessed by WebExtractor._preprocess_content() in src/web_extractor.py (line 281): BeautifulSoup with lxml parser strips 'script', 'style', 'header', 'footer', 'nav', 'aside' tags (using a frozen set at line 69), removes HTML comments and empty elements, then extracts cleaned text via soup.get_text(). Second, the clean text is sent to an LLM for semantic extraction guided by the prompt template in src/prompts.py (line 10). The prompt instructs the model to return pure JSON arrays when the user requests data export (detected via keyword matching on 'csv', 'json', 'excel', 'sql', 'html', 'export', 'extract', 'table'), or conversational text for general queries. There are no explicit CSS/XPath selectors — the LLM handles schema inference from context. For oversized pages, WebExtractor uses RecursiveCharacterTextSplitter with 32000-token chunks and 200-token overlap (lines 102-106), processes each chunk through the LLM independently, then merges JSON arrays via _merge_json_chunks() (line 412). Output formatting is handled by dedicated methods: _format_as_csv (line 442) writes to pandas DataFrame via csv.DictWriter; _format_as_excel (line 465) uses xlsxwriter; _format_as_sql (line 486) generates INSERT statements; _format_as_html (line 505) builds a table. There is no readability-style boilerplate removal beyond the tag stripping — the system relies on the LLM's comprehension to separate relevant content from noise.

How are LLMs used, if at all?

answered

LLMs are the core extraction engine, invoked after fetching. The WebExtractor in src/web_extractor.py (line 71) initializes a model via Models.get_model() in src/models.py, a factory that returns ChatOpenAI for GPT models (via LangChain), ChatGoogleGenerativeAI for Gemini, wraps ChatOpenAI pointed at a custom base URL for LiteLLM proxy models, or delegates to OllamaModel.generate() for local models. All models use the same unified prompt template from src/prompts.py (line 46). The prompt system has two modes: for data-export requests (detected by keywords: 'csv', 'json', 'excel', 'sql', 'html', 'export', 'extract', 'give me the data', 'table') the LLM is instructed to return only a valid JSON array; for conversational queries it responds in plain text with the configured netrunner AI persona. For large pages exceeding context windows, WebExtractor splits content via RecursiveCharacterTextSplitter (32000 tokens, 200 overlap), processes each chunk independently through the LLM, and merges JSON results via _merge_json_chunks(). Conversation history (last 10 messages, truncated to 500 chars each) is included in prompts for multi-turn context. A caching layer in _extract_info (line 310) caches LLM responses keyed by (content_hash, query, model_name), scoped only to export requests to avoid stale conversational responses. Cost controls are minimal: defaults to the cheapest model (gpt-4.1-mini), requires users to bring their own API keys, and provides no per-session token budget enforcement beyond the implicit chunking. Model classes are: GPT (gpt-4.1-mini, gpt-4o-mini, gpt-4, gpt-3.5-turbo, text-*), Gemini (gemini-1.5-flash, gemini-pro), Ollama (any installed model), LiteLLM (any model accessible via proxy).

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-bot evasion is handled primarily through Patchright (undetected Playwright fork) in src/scrapers/playwright_scraper.py. The ScraperConfig (line 36) exposes multiple toggles: use_stealth=True enables Patchright's automatic patches that remove webdriver flags, automation traces, and runtime.enable leaks; hide_webdriver=True ensures navigator.webdriver remains undefined; bypass_cloudflare=True reloads the page up to 3 times checking for Cloudflare challenge strings (line 716). The project deliberately avoids setting custom user_agent or viewport on browser contexts (line 474 comment) — those create detection signatures — and lets Patchright handle stealth automatically. A supplemental JavaScript init script at line 494 patches additional vectors: chrome.runtime, navigator.plugins, navigator.languages, navigator.connection.rtt. The simulate_human option (line 748) adds random scrolling, mouse movements, and element hovering. A 'Current Browser' mode (line 274) launches the user's real Chrome with remote debugging — using the actual browser profile for maximum undetectability. CAPTCHA handling at line 244 opens the page in non-headless mode, prints instructions to the console, and waits for user to press Enter after manual solving. Proxy support is flexible: Playwright accepts a per-request proxy string passed to browser launch options (line 438-439). Tor SOCKS5 at localhost:9050 is used for .onion routes via TorManager with randomized Tor Browser-like User-Agent rotation from a 3-agent pool (tor_config.py line 17-22). TorConfig supports auto_renew_circuit=True and verify_connection=True which checks against check.torproject.org/api/ip (tor_manager.py line 68-87). There is no IP rotation middleware, no proxy pool manager beyond Tor circuits, and no configurable request rate limiter — anti-bot measures focus on undetectability at the browser fingerprint level, not at the traffic-shaping level.

Editor's note. Correction: much of the listed stealth is not wired in. The init script calls are commented out, bypass_cloudflare() and simulate_human_behavior() are never called, auto_renew_circuit is unused, and WebExtractor always passes proxy=None for normal URLs; effective stealth is Patchright, real Chrome where available, no custom UA/viewport, and a small permissions.query patch.

How is crawling at scale implemented?

not applicable

The project does not implement web crawling at scale. There is no crawl queue, frontier manager, URL deduplication system, robots.txt parser, politeness delay configuration, or distributed worker architecture. Instead, it operates as a single-page or user-specified multi-page scraper within a Streamlit chat session. The closest feature is scrape_multiple_pages() in playwright_scraper.py (line 588), which takes a user-provided page range (e.g. '1-5'), constructs paginated URLs via pattern detection, and fetches them concurrently with a configurable semaphore (default 5 concurrent pages). But this is explicit per-session multi-page scraping, not autonomous crawling — the user specifies every URL and range. The application runs in a single Streamlit process with asyncio.run() wrapping async operations, no job queues, and no persistent crawl state. There is no robots.txt handling: compliance is left to the user's ethical judgment (mentioned only in the README).

What is the developer interface?

answered

The primary interface is a Streamlit web application launched via streamlit run main.py. The main.py file (line 420) constructs a chat-based UI with a sidebar for model selection (GPT, Gemini, LiteLLM, Ollama), service-status indicators, and chat-history management grouped by date. Users enter prompts that can be URLs to scrape or conversational queries about already-fetched content — extraction results are displayed as pandas DataFrames with download buttons for CSV/Excel and a 'Upload to Google Sheets' option. The Python library API centers on WebExtractor in src/web_extractor.py (line 71), which accepts a model name and optional ScraperConfig/TorConfig, then exposes process_query() for single-call scrape-and-extract. Lower-level: PlaywrightScraper (line 76 of playwright_scraper.py) for browser automation, TorScraper for onion sites, Models.get_model() in src/models.py for LLM instantiation. The StreamlitWebScraperChat class in app/streamlit_web_scraper_chat.py wraps WebExtractor and bridges Streamlit's synchronous event loop with the async pipeline via asyncio.run() (line 23). Output formats include JSON (markdown code blocks), CSV (in-UI DataFrame + download), Excel (via xlsxwriter bytes buffer), SQL (INSERT statements), and HTML tables — selected by user keywords in the query. There is no CLI interface (no argparse/click entry point), no REST service, no MCP server, and no HTTP API — the only invocation is streamlit run main.py. Language bindings: Python only. The Dockerfile (line 1 of Dockerfile) provides containerized deployment exposing ports 8501 (Streamlit UI), 9050 (Tor SOCKS5), and 9051 (Tor control), with an entrypoint that starts Tor and then launches the Streamlit app.