# vinyzu-archive/Botright

> Archived async Playwright wrapper with real-Chrome fingerprints, humanized input and local hCaptcha/reCAPTCHA solving; no extraction.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/vinyzu-archive/Botright (reviewed at commit `6e64fed40bb2d9dd93a5c2c57068994b702e884d`, 2026-09-09)
- Stars: 1036 · Language: Python · License: GPL-3.0
- Canonical page: https://llms-technical-reviews.com/p/botright/

## Overview

Botright is an async Python library that wraps Playwright to make an automated browser harder to detect. You replace the browser initialisation with `await botright.Botright()` and `await client.new_browser(proxy=...)`, and keep writing ordinary Playwright code. Underneath, it launches a Chromium-family browser already installed on your machine (not Playwright's bundled build), applies a fingerprint sampled from real Chrome data, aligns timezone and geolocation with the proxy's exit IP, moves the mouse along randomised Bézier curves, and adds `page.solve_hcaptcha()` and `page.solve_recaptcha()` methods backed by local vision models.

It is not a scraper in the extraction sense. There are no LLM calls, no HTML-to-Markdown conversion, no selectors-as-schema, and no crawler. You get the browser you would use *underneath* such a tool. The "AI" in its description is computer-vision CAPTCHA solving from two of the author's other packages, `hcaptcha_challenger` and `recognizer`.

Treat the project as frozen. The pinned repository lives under a `vinyzu-archive` organisation, the version is `0.5.1`, the last changelog entry is from January 2024, and the requirements pin `playwright==1.42.0`. The README marks hCaptcha solving as "outdated" and GeeTest as "Currently Not Available". The code agrees: `solve_geetest` raises `NotImplementedError`.

## Architecture

```mermaid
flowchart LR
  U["Your async code"] --> BR["Botright client"]
  BR --> ENG["get_browser_engine (local Chromium)"]
  BR --> NB["new_browser(proxy)"]
  NB --> PM["ProxyManager: IP + geo lookup"]
  NB --> FK["Faker: chrome-fingerprints"]
  NB --> CTX["launch_persistent_context"]
  CTX --> BC["BrowserContext wrapper"]
  BC --> PG["Page wrapper"]
  PG --> CDP["CDP UA + client hints override"]
  PG --> HM["Humanized Mouse / Keyboard"]
  PG --> HC["hCaptcha (hcaptcha_challenger)"]
  PG --> RC["reCAPTCHA (recognizer)"]
```

| Component | Path | Role |
|---|---|---|
| Client | `botright/botright.py` | `Botright` async object: Chrome flags, engine discovery, `new_browser`, cleanup |
| Proxy manager | `botright/modules/proxy_manager.py` | Parses proxy strings, checks the exit IP, looks up country, coordinates and timezone |
| Fingerprints | `botright/modules/faker.py` | Gets a fingerprint from `chrome-fingerprints` and a locale from the proxy country |
| Context wrapper | `botright/playwright_mock/browser.py` | Builds launch arguments, persistent context, image blocking, response cache |
| Page wrapper | `botright/playwright_mock/page.py` | CDP UA override, CAPTCHA methods, humanised `click`/`type` overrides |
| Input | `botright/playwright_mock/mouse.py`, `keyboard.py` | Bézier mouse paths and jittered key delays |
| hCaptcha | `botright/modules/hcaptcha.py` | Drives `hcaptcha_challenger.AgentT`, optional `rqdata` injection |
| Other wrappers | `botright/playwright_mock/` | `Locator`, `Frame`, `FrameLocator`, handles and routes re-wrapped so behaviour carries through |

## How a request flows

1. **Client init.** `Botright.__ainit__` first runs `hcaptcha_challenger.install(upgrade=True)` to fetch or refresh model assets. It then starts Playwright (or `undetected_playwright` when `use_undetected_playwright=True`), picks a browser, deletes leftover `botright-*` temp dirs, and builds a list of about 70 Chrome flags. These include `--disable-blink-features=AutomationControlled` and WebRTC `disable_non_proxied_udp`, plus `--fingerprinting-canvas-image-data-noise` when canvas spoofing is on ([botright.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L55-L132)).
2. **Engine choice.** `get_browser_engine` uses `pybrowsers` to find an installed browser. It tries `chromium` first (the author recommends Ungoogled Chromium), then `chrome`, then any non-Firefox/Edge/Opera/IE browser. It raises if none is found ([botright.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L191-L216)).
3. **Proxy check.** `new_browser` awaits `ProxyManager`, which accepts `ip:port`, `user:pass@ip:port` and colon-separated four-part forms. It asks five what-is-my-IP services for the exit IP through the proxy, then four geo-IP services for country, latitude, longitude and timezone. With no proxy, the same lookups describe your own IP ([proxy_manager.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/modules/proxy_manager.py#L37-L150)).
4. **Fingerprint.** `Faker` concurrently samples a `ChromeFingerprint` and maps the proxy country code to a locale. An unknown country raises `ValueError("Proxy Country not supported")` ([faker.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/modules/faker.py#L48-L315)).
5. **Launch.** `browser.new_browser` passes the fingerprint's user agent, screen and viewport, the proxy's timezone and geolocation, `color_scheme="dark"`, proxy credentials and `ignore_default_args=["--enable-automation"]` to `launch_persistent_context` with a fresh temp profile and the local browser's executable ([browser.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/browser.py#L31-L115)). It then grants notification and geolocation permissions ([botright.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L134-L166)).
6. **Page setup.** `new_page` opens a CDP session and sends both `Emulation.setUserAgentOverride` and `Network.setUserAgentOverride` with full client-hints metadata (brands, full version list, platform, architecture, bitness, model) from the fingerprint ([page.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/page.py#L162-L210)).
7. **Interaction.** `page.click(selector)` waits for the element, takes its bounding box, and calls `Mouse.click`. That moves along a humanised curve from the last cursor position, pauses 200 to 400 ms, then presses and releases ([page.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/page.py#L521-L564), [mouse.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/mouse.py#L165-L248)).

## Key components

### Playwright mock layer

Each wrapper subclasses the Playwright class, adopts the original's `_impl_obj`, keeps the original methods as `_origin_*`, and overrides the ones that return other Playwright objects. That way objects reached through `page.locator(...)` or `page.frames` are Botright types too. Route handlers are proxied so callbacks receive wrapped `Route`/`Request` objects. With `use_undetected_playwright`, methods that depend on the CDP `Runtime` domain (`expose_function`, `expose_binding`, `expect_console_message`) raise `NotSupportedError` ([browser.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/browser.py#L227-L281)). This works, but it is coupled tightly to Playwright internals, which is one reason the dependency is pinned.

### Human-like input

`HumanizeMouseTrajectory` (adapted from HumanCursor) picks two random knots within 80 px of the straight path, builds a Bézier curve, adds Gaussian distortion, and tweens to 100 points with an ease-out. Each point becomes one `mouse.move` ([mouse.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/mouse.py#L19-L45)). `Keyboard.type` sends one character at a time with 50 to 150 ms delays by default ([keyboard.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/keyboard.py#L13-L28)). Humanisation is always on. The `user_action_layer` flag only injects a full-screen canvas that draws coloured dots where the mouse moves and clicks. It is a debugging overlay, not a stealth feature.

### CAPTCHA solving

`hCaptcha.solve_hcaptcha` clicks the checkbox through `AgentT`, retries up to seven rounds (refreshing on a backcall), and returns the pass token. If you give it `rq_data`, it first intercepts `hcaptcha.com/getcaptcha/**` and re-posts the request with that `rqdata` ([hcaptcha.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/modules/hcaptcha.py#L29-L77)). On failure it returns an error *string* ("Exceeded maximum retry times…"), not `None`, so check the result before using it as a token. `solve_recaptcha` delegates to `recognizer`'s `AsyncChallenger`. `solve_geetest` raises `NotImplementedError`, even though a 368-line `geetest.py` and a TorchScript model still ship in the package ([page.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/page.py#L237-L260)).

### Bandwidth helpers

`block_images=True` aborts image requests by extension and by resource type. `cache_responses=True` serves documents, stylesheets, images, fonts and media from an in-memory dict keyed by URL, shared across all browsers of one client ([browser.py](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/browser.py#L175-L202)). Caching `document` responses means a second `goto` to the same URL replays stale HTML.

## Extending it

- **Existing Playwright code.** Everything after `new_browser()` is the Playwright async API, so extraction is up to you: `page.content()`, locators, or a downstream LLM extractor.
- **Launch overrides.** Extra keyword arguments to `new_browser` are merged over Botright's launch arguments, so you can override `locale`, `viewport` or `color_scheme`.
- **Stealth knobs.** `mask_fingerprint=False` keeps the real UA and screen (the README says one tough site only passes this way). `spoof_canvas`, `headless` and `use_undetected_playwright` are constructor flags.
- **Concurrency.** Not thread-safe. Use one client per thread, or several pages and browsers with `asyncio.gather`.

## Running it

- `pip install botright`, then `playwright install`. You also need a locally installed Chromium or Chrome; Ungoogled Chromium is recommended for canvas noise.
- Python 3.8+, async only. The default is a visible window (`headless=False`).
- The first run downloads `hcaptcha_challenger` assets, and every client start calls `install(upgrade=True)` again.
- Each `new_browser` call makes outbound requests to public IP and geo-IP services (ipify, ip-api.com, ipapi.co and others). Plan for those in locked-down networks.

## Strengths and caveats

- **Strength: coherent identity.** UA string, client hints, screen, timezone and geolocation all come from one fingerprint and the proxy's real exit location, rather than being randomised independently.
- **Strength: drop-in API.** Thorough wrapping means existing Playwright scripts mostly just work, with humanised input for free.
- **Strength: local CAPTCHA solving.** hCaptcha and reCAPTCHA image challenges are attempted with no paid solver API.
- **Caveat: abandoned and pinned.** The repository is archived, it is tied to Playwright 1.42 internals, GeeTest is disabled, and hCaptcha is flagged as outdated. Detection vendors have moved on since 2024, so the README's pass table is historical.
- **Caveat: locale inconsistency.** The browser locale and `Accept-Language` are hard-coded to `en-US`, whatever the proxy country. The computed per-country locale is never applied, and many of its codes are invalid anyway (`sp-ES`, `po-BR`, `ar-IL`). A German proxy therefore shows Berlin time with US English.
- **Caveat: small bugs.** `Locator.click` reads the bounding box *before* `scroll_into_view_if_needed` and clicks the stale coordinates. `solve_hcaptcha` returns an error string on failure. An unsupported proxy country aborts `new_browser`.
- **Caveat: not an extraction tool.** No LLM, content cleaning, crawling, rate limiting or proxy rotation. Pair it with something else for the data side.

*Sources: code at 6e64fed, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (9 pages), verified Q&A.*

## How vinyzu-archive/Botright answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

Botright uses **headless Chromium via Playwright** exclusively — there is no plain-HTTP fetch path. The `Botright.__ainit__` method (`botright/botright.py:55`) starts Playwright, then `get_browser_engine()` (`botright/botright.py:191-216`) locates a **locally-installed** Chromium-based browser (preferring Ungoogled Chromium, falling back to Chrome, then any Chromium variant). Browsers are launched as **persistent contexts** via `launch_persistent_context()` (`botright/playwright_mock/browser.py:88-95`), which preserves session state across pages.

**JS rendering** is handled automatically by Playwright's browser engine — there is no toggle. Pages are created through the `Page` wrapper (`botright/playwright_mock/page.py:58-74`), which calls `_mock_page()` to apply stealth patches via CDP (`page.py:162-210`), then returns the page.

The **waiting strategy** delegates to standard Playwright waits — `wait_for_selector`, `wait_for_timeout`. Custom interaction methods (`page.click`, `page.type`) add humanized mouse movement with Bézier curves and random delays, but these are for stealth rather than content-waiting.

**Content types**: PDF and image handling is left entirely to Playwright (e.g. `page.screenshot()`). Image **blocking** is available via the `block_images` flag (`botright/botright.py:36`), which installs a route interceptor aborting image resource types (`browser.py:191-202`). Response **caching** for static resources (document, stylesheet, image, media, font, manifest) is supported through `cache_responses` (`browser.py:175-189`), storing `APIResponse` objects in a dict keyed by URL.

There is no specialized PDF rendering or waiting beyond Playwright's built-in capabilities.

> **Editor's note.** Correction: the default is a headful browser (`headless=False`), launched from a locally installed Chromium/Chrome executable found via `pybrowsers`, not Playwright's bundled headless Chromium (botright/botright.py L55-L72, L191-L216).

Citations: [botright/botright.py:55-120](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L55-L120) · [botright/botright.py:191-216](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L191-L216) · [botright/playwright_mock/browser.py:88-115](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/browser.py#L88-L115) · [botright/playwright_mock/browser.py:175-202](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/browser.py#L175-L202) · [botright/playwright_mock/page.py:58-74](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/page.py#L58-L74)

### How is content extracted or converted? (not applicable)

Botright does **not** implement any content-extraction or conversion utilities. There is no HTML→Markdown conversion, no readability-style boilerplate removal, no CSS selector-based extraction schema, and no structured-output pipeline. The library exposes the full Playwright Page API (`page.content()`, `page.inner_text()`, `page.text_content()`, etc.) for users to extract whatever they need, but provides no helpers beyond that. Its purpose is undetected browser automation and CAPTCHA solving, not content parsing or transformation.



### How are LLMs used, if at all? (not applicable)

Botright does **not** use LLMs anywhere in its codebase. There are no prompts, no page-chunking strategies, no structured output schemas for LLM consumption, no cost controls, and no integration with any LLM provider (OpenAI, Anthropic, or otherwise). The 'AI' components mentioned are entirely **computer-vision models** used for CAPTCHA solving: `hcaptcha_challenger` uses vision models for hCaptcha image classification; `recognizer`/`reCognizer` handles reCaptcha; the commented-out GeeTest solver used CLIP-ViT-B-32 (a vision transformer) via `SentenceTransformer` for image similarity and YOLOv5 for object detection (`geetest_helpers.py:41,139`). These are image-classification and object-detection models, not language models. The `hcaptcha_challenger` model hub is installed via `solver.install(upgrade=True)` (`botright/botright.py:81`), but this downloads vision model weights only.


Citations: [botright/botright.py:81-81](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L81-L81) · [botright/modules/geetest_helpers.py:138-140](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/modules/geetest_helpers.py#L138-L140)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

Anti-bot evasion is Botright's raison d'être, implemented across multiple layers:

**Real browser.** Rather than Playwright's bundled browser, it finds a locally-installed Chromium via the `browsers` library, preferring **Ungoogled Chromium** for maximum stealth (`botright/botright.py:191-216`).

**Fingerprint spoofing.** `Faker.get_computer()` (`botright/modules/faker.py:48-52`) calls `chrome-fingerprints` (`AsyncFingerprintGenerator`) to generate a realistic fingerprint. The `Page._mock_page()` method then overrides the user agent via CDP using `Emulation.setUserAgentOverride` and `Network.setUserAgentOverride`, setting brands, fullVersionList, platform architecture, and model (`page.py:162-199`).

**CLI flags.** ~50 Chrome flags disable detection surfaces: `--disable-blink-features=AutomationControlled`, `--test-type`, `--incognito`, `--disable-extensions`, `--disable-ipc-flooding-protection`, `--hide-scrollbars`, `--safebrowsing-disable-auto-update`, and many more (`botright/botright.py:105-119`). `--ignore-default-args=["--enable-automation"]` is also set (`browser.py:64`).

**Canvas spoofing.** `spoof_canvas=True` triggers `--fingerprinting-canvas-image-data-noise` for Chromium or `--disable-reading-from-canvas` for others, foiling canvas-fingerprinting scripts (`botright/botright.py:126-131`).

**User-action layer.** When `user_action_layer=True`, a transparent overlay canvas is injected showing mouse movements as colored dots — mimicking human visual feedback (`page.py:201-210`).

**Humanized interaction.** The `Mouse` class (`mouse.py:18-248`) generates Bézier-curve mouse trajectories with random distortion and tweening, with random delays between actions. The `Keyboard` class (`keyboard.py:21-28`) types characters with per-character random delays.

**Proxy management.** `ProxyManager` (`proxy_manager.py:37-150`) parses proxy URLs, validates them with HTTP requests against geolocation APIs, resolves the proxy's timezone/location, and passes `http_credentials` and `browser_proxy` to the browser context.

**CAPTCHA solving.** hCaptcha, reCaptcha, and GeeTest (v3/v4 — temporarily unavailable in released code) are solved using local computer-vision models, requiring no external services. The README reports reCaptcha v3 scores of 0.9 and CreepJS scores around 65%.

Botright does **not** implement rate limiting, IP rotation scheduling, or robots.txt parsing.


Citations: [botright/botright.py:104-131](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L104-L131) · [botright/playwright_mock/page.py:162-210](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/page.py#L162-L210) · [botright/playwright_mock/mouse.py:18-248](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/mouse.py#L18-L248) · [botright/modules/faker.py:48-52](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/modules/faker.py#L48-L52) · [botright/modules/proxy_manager.py:37-150](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/modules/proxy_manager.py#L37-L150)

### How is crawling at scale implemented? (not applicable)

Botright does **not** implement any crawling-at-scale infrastructure. There are no URL queues, no URL deduplication, no crawl-depth limits, no robots.txt parsing, no politeness delays, and no distributed-worker coordination. The library provides `Botright` as a single-session entry point that manages one Playwright instance at a time; users who want multi-page crawling must build their own control flow on top of Botright's `new_page()` and `page.goto()` methods. The documentation explicitly notes the API is not thread-safe (`docs/index.rst:208-214`), recommending `asyncio.gather()` for concurrency.


Citations: [docs/index.rst:208-214](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/docs/index.rst#L208-L214)

### What is the developer interface? (answered)

Botright has a **single interface: an asynchronous Python library**. There is no CLI, REST service, MCP server, or graphical UI.

**Library API** — The developer imports `botright` and uses the async context manager pattern:

```python
botright_client = await botright.Botright(headless=False, block_images=True, cache_responses=False, ...)
browser = await botright_client.new_browser(proxy="user:pass@ip:port")
page = await browser.new_page()
await page.goto("https://example.com")
await page.solve_hcaptcha()  # custom method
await botright_client.close()
```

The `Botright` class (`botright/botright.py:27`) accepts config flags (headless, block_images, cache_responses, user_action_layer, scroll_into_view, spoof_canvas, mask_fingerprint, use_undetected_playwright). `new_browser()` (`botright/botright.py:134`) takes an optional `proxy` string plus any Playwright launch arguments, and returns a `BrowserContext` (`browser.py:118`).

**Playwright-compatible wrappers** — Every standard Playwright class is wrapped with Botright's overrides: `Page`, `BrowserContext`, `Locator`, `ElementHandle`, `Frame`, `FrameLocator`, `Mouse`, `Keyboard`, `Request`, `Response`, `Route` (re-exported from `botright/playwright_mock/__init__.py:1-12`). Users can write existing Playwright code by only changing the initialization.

**Extra Page methods** — `solve_hcaptcha()`, `get_hcaptcha()`, `solve_recaptcha()`, `solve_geetest()` are injected directly onto the `Page` class (`page.py:212-260`).

**Output formats** — None built in. Content is extracted via standard Playwright methods (HTML string via `page.content()`, text via locator access, screenshots via `page.screenshot()`).

**Language bindings** — Python 3.8+ only. Installed via `pip install botright` then `playwright install chromium`. Dependencies include Playwright, `hcaptcha_challenger`, `chrome-fingerprints`, and `recognizer` (for reCaptcha).


Citations: [botright/botright.py:27-91](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L27-L91) · [botright/botright.py:134-166](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/botright.py#L134-L166) · [botright/playwright_mock/__init__.py:1-12](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/__init__.py#L1-L12) · [botright/playwright_mock/page.py:212-260](https://github.com/vinyzu-archive/Botright/blob/6e64fed40bb2d9dd93a5c2c57068994b702e884d/botright/playwright_mock/page.py#L212-L260)
