# itsOwen/CyberScraper-2077

> Streamlit chat app that renders a page with Patchright, flattens it to text, and asks an LLM to answer or export it as JSON, CSV or SQL.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/itsOwen/CyberScraper-2077 (reviewed at commit `a260fa81f4a0c7c732885cceb68333c91fcda3d9`, 2026-09-27)
- Stars: 3280 · Language: Python · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/cyberscraper-2077/

## Overview

CyberScraper 2077 is a Streamlit chat app for one-off scraping. You paste a URL into the chat, it renders the page with Patchright (an undetectable fork of Playwright), strips the HTML to plain text, and keeps that text in memory. Each later message is answered by an LLM over that text. If your message contains words like "csv", "json", "excel", "sql" or "table", the prompt asks the model for a bare JSON array, and the app turns it into a DataFrame, a CSV or Excel download, SQL `INSERT`s, an HTML table, or a Google Sheet.

It is a personal tool with a themed UI, not a library or a service. There is no CLI, no HTTP API and no crawler. The model never sees HTML structure, selectors or a schema, only the visible text, so extraction quality is whatever the LLM can infer from flattened text. In return, the setup is small: pick a model (OpenAI, Gemini, Ollama, or anything behind a LiteLLM proxy), paste a URL, ask for a table.

Two extras stand out: `.onion` URLs are fetched through a local Tor SOCKS proxy, and a `-captcha` suffix opens a visible browser and waits for you to solve the challenge by hand.

## Architecture

```mermaid
flowchart LR
  U["User in Streamlit chat"] --> M["main.py UI"]
  M --> C["StreamlitWebScraperChat"]
  C --> X["WebExtractor.process_query"]
  X -->|"URL"| F["_fetch_url"]
  F -->|".onion"| T["TorScraper via SOCKS 9050"]
  F -->|"other"| P["PlaywrightScraper (Patchright)"]
  P --> PRE["_preprocess_content to text"]
  T --> PRE
  X -->|"question"| E["_extract_info"]
  E --> L["LLM: OpenAI, Gemini, Ollama, LiteLLM"]
  L --> FMT["_format_result"]
  FMT --> OUT["JSON, CSV, Excel, SQL, HTML, Sheets"]
```

| Component | Path | Role |
|---|---|---|
| Streamlit app | `main.py` | Sidebar model picker, chat history in `chat_history.json`, result rendering and downloads |
| Sync bridge | `app/streamlit_web_scraper_chat.py` | Wraps `WebExtractor` and runs each message with `asyncio.run` |
| Orchestrator | `src/web_extractor.py` | URL detection, fetch, preprocessing, LLM calls, chunking, cache, output formatting |
| Browser scraper | `src/scrapers/playwright_scraper.py` | `ScraperConfig`, pooled Patchright browser, multi-page and CAPTCHA modes |
| Tor scraper | `src/scrapers/tor/` | aiohttp over SOCKS5 for `.onion` hosts, Tor connectivity check |
| Model factory | `src/models.py`, `src/ollama_models.py` | LangChain chat models, LiteLLM proxy via `ChatOpenAI`, Ollama over its REST API |
| Prompt | `src/prompts.py` | One shared template with a data-export mode and a conversational mode |
| Sheets export | `src/utils/google_sheets_utils.py` | OAuth flow and upload of a DataFrame |

## How a request flows

Take two messages: `https://example.com/products?page=1 1-3`, then `give me a csv of name and price`.

1. **Bridge.** `StreamlitWebScraperChat.process_message` runs `WebExtractor.process_query` inside `asyncio.run`, with a Streamlit placeholder as the progress callback ([streamlit_web_scraper_chat.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/app/streamlit_web_scraper_chat.py#L7-L23)). `main.py` builds the `ScraperConfig` from the sidebar, with `delay_after_load=5` ([main.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/main.py#L178-L194)).
2. **Parse the message.** `process_query` finds the first URL with a regex. The token after it is treated as a page spec if it contains only digits, commas and dashes, the next token as an optional URL pattern, and `-captcha` anywhere turns on CAPTCHA mode ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L191-L222)).
3. **Fetch.** `_fetch_url` sends `.onion` hosts to `TorScraper` and everything else to `PlaywrightScraper.fetch_content` with `proxy=None` ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L224-L279)). With pages, `detect_url_pattern` finds the first numeric query parameter (`page={page}`) or path segment, and `scrape_multiple_pages` opens one fresh context per page under a semaphore of 5, each followed by a random 0.5 to 1.5 s pause ([playwright_scraper.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L588-L684), [L776-L851](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L776-L851)).
4. **Flatten.** The HTML of all pages is joined and passed through `_preprocess_content`: drop `script`, `style`, `header`, `footer`, `nav` and `aside`, drop comments and empty tags, then `get_text()` and collapse whitespace ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L281-L302)). A content hash resets the query cache when the page changes.
5. **Ask.** The second message has no URL, so `_extract_info` runs. If the text fits in `max_tokens - 1000` (128k for `gpt-4.1-mini` and `gpt-4o-mini`, 16,385 otherwise), it makes one call; otherwise it splits into 32,000-token chunks, calls the model per chunk and merges the JSON arrays ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L304-L348)). `_call_model` fills the shared prompt with the last 10 chat turns (each cut to 500 characters), the page text and the query ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L121-L160)).
6. **Format.** `_format_result` keyword-matches the query again, pulls a JSON array out of the reply (direct parse, fenced block, or regex), and for "csv" returns a CSV string plus a DataFrame ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L350-L407)). `main.py` shows the table and download buttons.

## Key components

### Prompt and output modes

There is one prompt for every model. It sets a "netrunner" persona, says to return only a JSON array for export-style requests, to use `"N/A"` for missing fields and never to invent data, and to answer in plain text otherwise ([prompts.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/prompts.py#L9-L50)). Mode selection is plain substring matching on the user's message, so "extract" or "table" anywhere in a question switches to JSON. The SQL formatter builds a `CREATE TABLE` with all columns as `TEXT` and escapes single quotes ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L486-L503)).

### Model routing

`WebExtractor.__init__` picks the backend from the model name: `ollama:` goes to `OllamaModel`, `gemini-` to `ChatGoogleGenerativeAI`, everything else to `Models.get_model` ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L71-L111)). The factory accepts four OpenAI chat names, `text-*` completion models, and `litellm:<model>`, which becomes a `ChatOpenAI` pointed at `LITELLM_BASE_URL` ([models.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/models.py#L11-L54)). Ollama is called directly at `/api/generate` with streaming, opening a fresh aiohttp session per call because each Streamlit message runs in a new event loop ([ollama_models.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/ollama_models.py#L21-L48)).

### Browser and stealth

`ScraperConfig` exposes many toggles, but the effective stealth is Patchright itself plus real Chrome where available (`channel="chrome"` except on ARM64 Linux), no custom user agent or viewport, and one small `permissions.query` patch per page ([playwright_scraper.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L26-L74), [L420-L482](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L420-L482), [L539-L567](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L539-L567)). The extra init script is defined but its calls are commented out. `bypass_cloudflare()` and `simulate_human_behavior()` exist but nothing calls them, and the `bypass_cloudflare` flag is never read ([playwright_scraper.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L716-L774)). The "Use Current Browser" option launches your installed Chrome with a temp profile and `--remote-debugging-port=9222`, then connects over CDP ([playwright_scraper.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L274-L312)).

### CAPTCHA mode

With `-captcha`, the browser launches headed, navigates, and blocks on `aioconsole.ainput()` until you press Enter in the terminal that runs Streamlit, then scrapes the current page and any further pages in the same tab, and closes the browser ([playwright_scraper.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L189-L272)). That works on a desktop, not in Docker.

### Tor

`TorManager` opens one aiohttp session through `socks5://127.0.0.1:9050` with headers chosen once from three Firefox user agents, checks `check.torproject.org/api/ip` before each fetch, and returns the raw response text ([tor_manager.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/tor/tor_manager.py#L18-L113)). There is no JavaScript rendering on this path, and `auto_renew_circuit` in `TorConfig` is never used ([tor_config.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/tor/tor_config.py#L4-L22)).

## Extending it

- **New model provider.** Add a `case` in `Models.get_model` and a matching prefix in `get_prompt_for_model`, or put the provider behind a LiteLLM proxy and use `litellm:<name>` with no code change.
- **New fetcher.** Subclass `BaseScraper` (async `fetch_content` and `extract`) and route to it in `WebExtractor._fetch_url`, the same way `.onion` URLs go to `TorScraper`.
- **Use it without Streamlit.** `WebExtractor(model_name=...)` plus `await process_query(url)` and then `await process_query("give me json of ...")` works as a two-step library call, since it holds the page text as instance state.
- **Better extraction.** The quickest win is to stop flattening to text: pass cleaned HTML or markdown into `_preprocess_content`, or add a schema to the prompt.

## Running it

- **Local.** Python 3.10+ (the code uses `match`), `pip install -r requirements.txt`, `patchright install chromium` (or `chrome`), set `OPENAI_API_KEY` or `GOOGLE_API_KEY`, then `streamlit run main.py`. Ollama is reached at `OLLAMA_BASE_URL` (default `http://localhost:11434`).
- **Docker.** The image is Python 3.12 slim with Tor and Patchright Chromium; a generated `run.sh` starts Tor, waits for port 9050, and launches Streamlit on 8501 ([Dockerfile](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/Dockerfile#L1-L113)). "Use Current Browser" and CAPTCHA mode do not work inside the container.

## Strengths and caveats

- **Strength: low ceremony.** Paste a URL, name the fields, download a CSV. Pagination by `1-5` and `.onion` support are built in.
- **Strength: model freedom.** OpenAI, Gemini, any Ollama model, and anything behind a LiteLLM proxy share one prompt and one code path.
- **Strength: honest stealth defaults.** It relies on Patchright and real Chrome instead of fragile JS spoofing, and leaves user agent and viewport alone.
- **Caveat: text-only context.** Links, attributes, image URLs and table structure are lost before the LLM sees the page, and `header`, `nav` and `footer` content is always dropped.
- **Caveat: brittle large-page path.** Chunked results are merged with a plain `json.loads` on each reply, so fenced or chatty replies are silently dropped, and a conversational question over a long page comes back as a JSON array or as `[]` ([web_extractor.py](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L412-L423)).
- **Caveat: advertised anti-bot features are idle.** Cloudflare retry, human simulation, the init script and Tor circuit renewal are not wired in, and the app always passes `proxy=None`, so there is no proxy support for normal sites.
- **Caveat: single-user app.** Chat history is a local `chat_history.json`, each message runs in its own event loop, and CAPTCHA mode waits on the server's terminal.

*Sources: code at a260fa8, deepwiki-open wiki (10 pages), verified Q&A.*

## How itsOwen/CyberScraper-2077 answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

Pages are fetched primarily via Patchright (undetected Playwright fork) for full headless browser automation with JavaScript rendering. The `PlaywrightScraper` class in `src/scrapers/playwright_scraper.py` orchestrates this: it launches a Chromium browser (or real Chrome on x86 platforms), creates browser contexts, navigates via `page.goto()` with a configurable wait strategy (default `domcontentloaded` in `ScraperConfig`, line 50), and extracts raw HTML via `page.content()`. A 'Current Browser' mode at line 274 launches the user's real Chrome with remote debugging on port 9222 and connects over CDP, using the user's own profile for maximum undetectability. For `.onion` URLs, `TorScraper` in `src/scrapers/tor/tor_scraper.py` routes via SOCKS5 at `127.0.0.1:9050` using `aiohttp_socks.ProxyConnector` (tor_manager.py line 53-54). Multi-page fetching supports URL pattern detection in query parameters or path segments (lines 776-802) and concurrent scraping via `asyncio.gather` with a configurable semaphore cap (`max_concurrent_pages`, line 650). A shared aiohttp `ClientSession` singleton in `src/http_client.py` provides plain-HTTP connection-pooled fetching for non-browser use. There is no PDF rendering, no image content extraction — the scraper extracts only HTML content from the browser.

> **Editor's note.** Correction: src/http_client.py is not imported anywhere, so there is no plain-HTTP fetch path; non-onion pages always go through Patchright and onion pages through TorManager's own aiohttp session.

Citations: [src/scrapers/playwright_scraper.py:36-73](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L36-L73) · [src/scrapers/playwright_scraper.py:244-272](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L244-L272) · [src/scrapers/playwright_scraper.py:586-684](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L586-L684) · [src/scrapers/playwright_scraper.py:686-714](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L686-L714) · [src/scrapers/tor/tor_manager.py:49-60](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/tor/tor_manager.py#L49-L60) · [src/http_client.py:29-52](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/http_client.py#L29-L52)

### How is content extracted or converted? (answered)

Extraction is a two-stage pipeline. First, raw HTML is preprocessed by `WebExtractor._preprocess_content()` in `src/web_extractor.py` (line 281): BeautifulSoup with lxml parser strips 'script', 'style', 'header', 'footer', 'nav', 'aside' tags (using a frozen set at line 69), removes HTML comments and empty elements, then extracts cleaned text via `soup.get_text()`. Second, the clean text is sent to an LLM for semantic extraction guided by the prompt template in `src/prompts.py` (line 10). The prompt instructs the model to return pure JSON arrays when the user requests data export (detected via keyword matching on 'csv', 'json', 'excel', 'sql', 'html', 'export', 'extract', 'table'), or conversational text for general queries. There are no explicit CSS/XPath selectors — the LLM handles schema inference from context. For oversized pages, `WebExtractor` uses `RecursiveCharacterTextSplitter` with 32000-token chunks and 200-token overlap (lines 102-106), processes each chunk through the LLM independently, then merges JSON arrays via `_merge_json_chunks()` (line 412). Output formatting is handled by dedicated methods: `_format_as_csv` (line 442) writes to pandas DataFrame via `csv.DictWriter`; `_format_as_excel` (line 465) uses xlsxwriter; `_format_as_sql` (line 486) generates INSERT statements; `_format_as_html` (line 505) builds a table. There is no readability-style boilerplate removal beyond the tag stripping — the system relies on the LLM's comprehension to separate relevant content from noise.


Citations: [src/web_extractor.py:281-302](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L281-L302) · [src/web_extractor.py:102-106](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L102-L106) · [src/web_extractor.py:304-348](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L304-L348) · [src/web_extractor.py:412-423](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L412-L423) · [src/web_extractor.py:442-463](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L442-L463) · [src/prompts.py:9-44](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/prompts.py#L9-L44)

### How are LLMs used, if at all? (answered)

LLMs are the core extraction engine, invoked after fetching. The `WebExtractor` in `src/web_extractor.py` (line 71) initializes a model via `Models.get_model()` in `src/models.py`, a factory that returns `ChatOpenAI` for GPT models (via LangChain), `ChatGoogleGenerativeAI` for Gemini, wraps `ChatOpenAI` pointed at a custom base URL for LiteLLM proxy models, or delegates to `OllamaModel.generate()` for local models. All models use the same unified prompt template from `src/prompts.py` (line 46). The prompt system has two modes: for data-export requests (detected by keywords: 'csv', 'json', 'excel', 'sql', 'html', 'export', 'extract', 'give me the data', 'table') the LLM is instructed to return only a valid JSON array; for conversational queries it responds in plain text with the configured netrunner AI persona. For large pages exceeding context windows, `WebExtractor` splits content via `RecursiveCharacterTextSplitter` (32000 tokens, 200 overlap), processes each chunk independently through the LLM, and merges JSON results via `_merge_json_chunks()`. Conversation history (last 10 messages, truncated to 500 chars each) is included in prompts for multi-turn context. A caching layer in `_extract_info` (line 310) caches LLM responses keyed by (content_hash, query, model_name), scoped only to export requests to avoid stale conversational responses. Cost controls are minimal: defaults to the cheapest model (`gpt-4.1-mini`), requires users to bring their own API keys, and provides no per-session token budget enforcement beyond the implicit chunking. Model classes are: GPT (gpt-4.1-mini, gpt-4o-mini, gpt-4, gpt-3.5-turbo, text-*), Gemini (gemini-1.5-flash, gemini-pro), Ollama (any installed model), LiteLLM (any model accessible via proxy).


Citations: [src/models.py:11-54](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/models.py#L11-L54) · [src/ollama_models.py:16-61](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/ollama_models.py#L16-L61) · [src/prompts.py:9-50](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/prompts.py#L9-L50) · [src/web_extractor.py:139-160](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L139-L160) · [src/web_extractor.py:304-348](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L304-L348) · [src/web_extractor.py:102-106](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L102-L106)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

Anti-bot evasion is handled primarily through Patchright (undetected Playwright fork) in `src/scrapers/playwright_scraper.py`. The `ScraperConfig` (line 36) exposes multiple toggles: `use_stealth=True` enables Patchright's automatic patches that remove webdriver flags, automation traces, and runtime.enable leaks; `hide_webdriver=True` ensures navigator.webdriver remains undefined; `bypass_cloudflare=True` reloads the page up to 3 times checking for Cloudflare challenge strings (line 716). The project deliberately avoids setting custom `user_agent` or `viewport` on browser contexts (line 474 comment) — those create detection signatures — and lets Patchright handle stealth automatically. A supplemental JavaScript init script at line 494 patches additional vectors: `chrome.runtime`, `navigator.plugins`, `navigator.languages`, `navigator.connection.rtt`. The `simulate_human` option (line 748) adds random scrolling, mouse movements, and element hovering. A 'Current Browser' mode (line 274) launches the user's real Chrome with remote debugging — using the actual browser profile for maximum undetectability. CAPTCHA handling at line 244 opens the page in non-headless mode, prints instructions to the console, and waits for user to press Enter after manual solving. Proxy support is flexible: Playwright accepts a per-request proxy string passed to browser launch options (line 438-439). Tor SOCKS5 at `localhost:9050` is used for `.onion` routes via `TorManager` with randomized Tor Browser-like User-Agent rotation from a 3-agent pool (tor_config.py line 17-22). `TorConfig` supports `auto_renew_circuit=True` and `verify_connection=True` which checks against `check.torproject.org/api/ip` (tor_manager.py line 68-87). There is no IP rotation middleware, no proxy pool manager beyond Tor circuits, and no configurable request rate limiter — anti-bot measures focus on undetectability at the browser fingerprint level, not at the traffic-shaping level.

> **Editor's note.** Correction: much of the listed stealth is not wired in. The init script calls are commented out, bypass_cloudflare() and simulate_human_behavior() are never called, auto_renew_circuit is unused, and WebExtractor always passes proxy=None for normal URLs; effective stealth is Patchright, real Chrome where available, no custom UA/viewport, and a small permissions.query patch.

Citations: [src/scrapers/playwright_scraper.py:36-73](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L36-L73) · [src/scrapers/playwright_scraper.py:484-537](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L484-L537) · [src/scrapers/playwright_scraper.py:539-567](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L539-L567) · [src/scrapers/playwright_scraper.py:716-746](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L716-L746) · [src/scrapers/playwright_scraper.py:748-774](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L748-L774) · [src/scrapers/tor/tor_manager.py:34-87](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/tor/tor_manager.py#L34-L87)

### How is crawling at scale implemented? (not applicable)

The project does not implement web crawling at scale. There is no crawl queue, frontier manager, URL deduplication system, robots.txt parser, politeness delay configuration, or distributed worker architecture. Instead, it operates as a single-page or user-specified multi-page scraper within a Streamlit chat session. The closest feature is `scrape_multiple_pages()` in `playwright_scraper.py` (line 588), which takes a user-provided page range (e.g. '1-5'), constructs paginated URLs via pattern detection, and fetches them concurrently with a configurable semaphore (default 5 concurrent pages). But this is explicit per-session multi-page scraping, not autonomous crawling — the user specifies every URL and range. The application runs in a single Streamlit process with `asyncio.run()` wrapping async operations, no job queues, and no persistent crawl state. There is no robots.txt handling: compliance is left to the user's ethical judgment (mentioned only in the README).


Citations: [src/scrapers/playwright_scraper.py:588-684](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L588-L684) · [src/web_extractor.py:224-258](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L224-L258) · [src/scrapers/playwright_scraper.py:776-826](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/scrapers/playwright_scraper.py#L776-L826) · [README.md:259-265](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/README.md#L259-L265)

### What is the developer interface? (answered)

The primary interface is a **Streamlit web application** launched via `streamlit run main.py`. The `main.py` file (line 420) constructs a chat-based UI with a sidebar for model selection (GPT, Gemini, LiteLLM, Ollama), service-status indicators, and chat-history management grouped by date. Users enter prompts that can be URLs to scrape or conversational queries about already-fetched content — extraction results are displayed as pandas DataFrames with download buttons for CSV/Excel and a 'Upload to Google Sheets' option. The Python **library API** centers on `WebExtractor` in `src/web_extractor.py` (line 71), which accepts a model name and optional `ScraperConfig`/`TorConfig`, then exposes `process_query()` for single-call scrape-and-extract. Lower-level: `PlaywrightScraper` (line 76 of `playwright_scraper.py`) for browser automation, `TorScraper` for onion sites, `Models.get_model()` in `src/models.py` for LLM instantiation. The `StreamlitWebScraperChat` class in `app/streamlit_web_scraper_chat.py` wraps `WebExtractor` and bridges Streamlit's synchronous event loop with the async pipeline via `asyncio.run()` (line 23). Output formats include JSON (markdown code blocks), CSV (in-UI DataFrame + download), Excel (via xlsxwriter bytes buffer), SQL (INSERT statements), and HTML tables — selected by user keywords in the query. There is **no CLI interface** (no `argparse`/`click` entry point), **no REST service**, **no MCP server**, and **no HTTP API** — the only invocation is `streamlit run main.py`. Language bindings: **Python only**. The Dockerfile (line 1 of `Dockerfile`) provides containerized deployment exposing ports 8501 (Streamlit UI), 9050 (Tor SOCKS5), and 9051 (Tor control), with an entrypoint that starts Tor and then launches the Streamlit app.


Citations: [main.py:420-613](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/main.py#L420-L613) · [app/streamlit_web_scraper_chat.py:1-23](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/app/streamlit_web_scraper_chat.py#L1-L23) · [src/web_extractor.py:71-111](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L71-L111) · [src/web_extractor.py:375-407](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/src/web_extractor.py#L375-L407) · [main.py:454-475](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/main.py#L454-L475) · [Dockerfile:1-114](https://github.com/itsOwen/CyberScraper-2077/blob/a260fa81f4a0c7c732885cceb68333c91fcda3d9/Dockerfile#L1-L114)
