LLMs Technical Reviews
Home / AI web scraping / Botright

vinyzu-archive/Botright

Archived async Playwright wrapper with real-Chrome fingerprints, humanized input and local hCaptcha/reCAPTCHA solving; no extraction.

GitHub ↗★ 1.0kPythonGPL-3.0commit 6e64fed · 2026-09-09

Overview

Botright is an async Python library that wraps Playwright to make an automated browser harder to detect. You replace the browser initialisation with await botright.Botright() and await client.new_browser(proxy=...), and keep writing ordinary Playwright code. Underneath, it launches a Chromium-family browser already installed on your machine (not Playwright’s bundled build), applies a fingerprint sampled from real Chrome data, aligns timezone and geolocation with the proxy’s exit IP, moves the mouse along randomised Bézier curves, and adds page.solve_hcaptcha() and page.solve_recaptcha() methods backed by local vision models.

It is not a scraper in the extraction sense. There are no LLM calls, no HTML-to-Markdown conversion, no selectors-as-schema, and no crawler. You get the browser you would use underneath such a tool. The “AI” in its description is computer-vision CAPTCHA solving from two of the author’s other packages, hcaptcha_challenger and recognizer.

Treat the project as frozen. The pinned repository lives under a vinyzu-archive organisation, the version is 0.5.1, the last changelog entry is from January 2024, and the requirements pin playwright==1.42.0. The README marks hCaptcha solving as “outdated” and GeeTest as “Currently Not Available”. The code agrees: solve_geetest raises NotImplementedError.

Architecture

flowchart LR
  U["Your async code"] --> BR["Botright client"]
  BR --> ENG["get_browser_engine (local Chromium)"]
  BR --> NB["new_browser(proxy)"]
  NB --> PM["ProxyManager: IP + geo lookup"]
  NB --> FK["Faker: chrome-fingerprints"]
  NB --> CTX["launch_persistent_context"]
  CTX --> BC["BrowserContext wrapper"]
  BC --> PG["Page wrapper"]
  PG --> CDP["CDP UA + client hints override"]
  PG --> HM["Humanized Mouse / Keyboard"]
  PG --> HC["hCaptcha (hcaptcha_challenger)"]
  PG --> RC["reCAPTCHA (recognizer)"]
Component Path Role
Client botright/botright.py Botright async object: Chrome flags, engine discovery, new_browser, cleanup
Proxy manager botright/modules/proxy_manager.py Parses proxy strings, checks the exit IP, looks up country, coordinates and timezone
Fingerprints botright/modules/faker.py Gets a fingerprint from chrome-fingerprints and a locale from the proxy country
Context wrapper botright/playwright_mock/browser.py Builds launch arguments, persistent context, image blocking, response cache
Page wrapper botright/playwright_mock/page.py CDP UA override, CAPTCHA methods, humanised click/type overrides
Input botright/playwright_mock/mouse.py, keyboard.py Bézier mouse paths and jittered key delays
hCaptcha botright/modules/hcaptcha.py Drives hcaptcha_challenger.AgentT, optional rqdata injection
Other wrappers botright/playwright_mock/ Locator, Frame, FrameLocator, handles and routes re-wrapped so behaviour carries through

How a request flows

  1. Client init. Botright.__ainit__ first runs hcaptcha_challenger.install(upgrade=True) to fetch or refresh model assets. It then starts Playwright (or undetected_playwright when use_undetected_playwright=True), picks a browser, deletes leftover botright-* temp dirs, and builds a list of about 70 Chrome flags. These include --disable-blink-features=AutomationControlled and WebRTC disable_non_proxied_udp, plus --fingerprinting-canvas-image-data-noise when canvas spoofing is on (botright.py).
  2. Engine choice. get_browser_engine uses pybrowsers to find an installed browser. It tries chromium first (the author recommends Ungoogled Chromium), then chrome, then any non-Firefox/Edge/Opera/IE browser. It raises if none is found (botright.py).
  3. Proxy check. new_browser awaits ProxyManager, which accepts ip:port, user:pass@ip:port and colon-separated four-part forms. It asks five what-is-my-IP services for the exit IP through the proxy, then four geo-IP services for country, latitude, longitude and timezone. With no proxy, the same lookups describe your own IP (proxy_manager.py).
  4. Fingerprint. Faker concurrently samples a ChromeFingerprint and maps the proxy country code to a locale. An unknown country raises ValueError("Proxy Country not supported") (faker.py).
  5. Launch. browser.new_browser passes the fingerprint’s user agent, screen and viewport, the proxy’s timezone and geolocation, color_scheme="dark", proxy credentials and ignore_default_args=["--enable-automation"] to launch_persistent_context with a fresh temp profile and the local browser’s executable (browser.py). It then grants notification and geolocation permissions (botright.py).
  6. Page setup. new_page opens a CDP session and sends both Emulation.setUserAgentOverride and Network.setUserAgentOverride with full client-hints metadata (brands, full version list, platform, architecture, bitness, model) from the fingerprint (page.py).
  7. Interaction. page.click(selector) waits for the element, takes its bounding box, and calls Mouse.click. That moves along a humanised curve from the last cursor position, pauses 200 to 400 ms, then presses and releases (page.py, mouse.py).

Key components

Playwright mock layer

Each wrapper subclasses the Playwright class, adopts the original’s _impl_obj, keeps the original methods as _origin_*, and overrides the ones that return other Playwright objects. That way objects reached through page.locator(...) or page.frames are Botright types too. Route handlers are proxied so callbacks receive wrapped Route/Request objects. With use_undetected_playwright, methods that depend on the CDP Runtime domain (expose_function, expose_binding, expect_console_message) raise NotSupportedError (browser.py). This works, but it is coupled tightly to Playwright internals, which is one reason the dependency is pinned.

Human-like input

HumanizeMouseTrajectory (adapted from HumanCursor) picks two random knots within 80 px of the straight path, builds a Bézier curve, adds Gaussian distortion, and tweens to 100 points with an ease-out. Each point becomes one mouse.move (mouse.py). Keyboard.type sends one character at a time with 50 to 150 ms delays by default (keyboard.py). Humanisation is always on. The user_action_layer flag only injects a full-screen canvas that draws coloured dots where the mouse moves and clicks. It is a debugging overlay, not a stealth feature.

CAPTCHA solving

hCaptcha.solve_hcaptcha clicks the checkbox through AgentT, retries up to seven rounds (refreshing on a backcall), and returns the pass token. If you give it rq_data, it first intercepts hcaptcha.com/getcaptcha/** and re-posts the request with that rqdata (hcaptcha.py). On failure it returns an error string (“Exceeded maximum retry times…”), not None, so check the result before using it as a token. solve_recaptcha delegates to recognizer’s AsyncChallenger. solve_geetest raises NotImplementedError, even though a 368-line geetest.py and a TorchScript model still ship in the package (page.py).

Bandwidth helpers

block_images=True aborts image requests by extension and by resource type. cache_responses=True serves documents, stylesheets, images, fonts and media from an in-memory dict keyed by URL, shared across all browsers of one client (browser.py). Caching document responses means a second goto to the same URL replays stale HTML.

Extending it

  • Existing Playwright code. Everything after new_browser() is the Playwright async API, so extraction is up to you: page.content(), locators, or a downstream LLM extractor.
  • Launch overrides. Extra keyword arguments to new_browser are merged over Botright’s launch arguments, so you can override locale, viewport or color_scheme.
  • Stealth knobs. mask_fingerprint=False keeps the real UA and screen (the README says one tough site only passes this way). spoof_canvas, headless and use_undetected_playwright are constructor flags.
  • Concurrency. Not thread-safe. Use one client per thread, or several pages and browsers with asyncio.gather.

Running it

  • pip install botright, then playwright install. You also need a locally installed Chromium or Chrome; Ungoogled Chromium is recommended for canvas noise.
  • Python 3.8+, async only. The default is a visible window (headless=False).
  • The first run downloads hcaptcha_challenger assets, and every client start calls install(upgrade=True) again.
  • Each new_browser call makes outbound requests to public IP and geo-IP services (ipify, ip-api.com, ipapi.co and others). Plan for those in locked-down networks.

Strengths and caveats

  • Strength: coherent identity. UA string, client hints, screen, timezone and geolocation all come from one fingerprint and the proxy’s real exit location, rather than being randomised independently.
  • Strength: drop-in API. Thorough wrapping means existing Playwright scripts mostly just work, with humanised input for free.
  • Strength: local CAPTCHA solving. hCaptcha and reCAPTCHA image challenges are attempted with no paid solver API.
  • Caveat: abandoned and pinned. The repository is archived, it is tied to Playwright 1.42 internals, GeeTest is disabled, and hCaptcha is flagged as outdated. Detection vendors have moved on since 2024, so the README’s pass table is historical.
  • Caveat: locale inconsistency. The browser locale and Accept-Language are hard-coded to en-US, whatever the proxy country. The computed per-country locale is never applied, and many of its codes are invalid anyway (sp-ES, po-BR, ar-IL). A German proxy therefore shows Berlin time with US English.
  • Caveat: small bugs. Locator.click reads the bounding box before scroll_into_view_if_needed and clicks the stale coordinates. solve_hcaptcha returns an error string on failure. An unsupported proxy country aborts new_browser.
  • Caveat: not an extraction tool. No LLM, content cleaning, crawling, rate limiting or proxy rotation. Pair it with something else for the data side.

Sources: code at 6e64fed, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (9 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 6e64fed. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Botright uses headless Chromium via Playwright exclusively — there is no plain-HTTP fetch path. The Botright.__ainit__ method (botright/botright.py:55) starts Playwright, then get_browser_engine() (botright/botright.py:191-216) locates a locally-installed Chromium-based browser (preferring Ungoogled Chromium, falling back to Chrome, then any Chromium variant). Browsers are launched as persistent contexts via launch_persistent_context() (botright/playwright_mock/browser.py:88-95), which preserves session state across pages.

JS rendering is handled automatically by Playwright's browser engine — there is no toggle. Pages are created through the Page wrapper (botright/playwright_mock/page.py:58-74), which calls _mock_page() to apply stealth patches via CDP (page.py:162-210), then returns the page.

The waiting strategy delegates to standard Playwright waits — wait_for_selector, wait_for_timeout. Custom interaction methods (page.click, page.type) add humanized mouse movement with Bézier curves and random delays, but these are for stealth rather than content-waiting.

Content types: PDF and image handling is left entirely to Playwright (e.g. page.screenshot()). Image blocking is available via the block_images flag (botright/botright.py:36), which installs a route interceptor aborting image resource types (browser.py:191-202). Response caching for static resources (document, stylesheet, image, media, font, manifest) is supported through cache_responses (browser.py:175-189), storing APIResponse objects in a dict keyed by URL.

There is no specialized PDF rendering or waiting beyond Playwright's built-in capabilities.

Editor's note. Correction: the default is a headful browser (headless=False), launched from a locally installed Chromium/Chrome executable found via pybrowsers, not Playwright's bundled headless Chromium (botright/botright.py L55-L72, L191-L216).

How is content extracted or converted?

not applicable

Botright does not implement any content-extraction or conversion utilities. There is no HTML→Markdown conversion, no readability-style boilerplate removal, no CSS selector-based extraction schema, and no structured-output pipeline. The library exposes the full Playwright Page API (page.content(), page.inner_text(), page.text_content(), etc.) for users to extract whatever they need, but provides no helpers beyond that. Its purpose is undetected browser automation and CAPTCHA solving, not content parsing or transformation.

How are LLMs used, if at all?

not applicable

Botright does not use LLMs anywhere in its codebase. There are no prompts, no page-chunking strategies, no structured output schemas for LLM consumption, no cost controls, and no integration with any LLM provider (OpenAI, Anthropic, or otherwise). The 'AI' components mentioned are entirely computer-vision models used for CAPTCHA solving: hcaptcha_challenger uses vision models for hCaptcha image classification; recognizer/reCognizer handles reCaptcha; the commented-out GeeTest solver used CLIP-ViT-B-32 (a vision transformer) via SentenceTransformer for image similarity and YOLOv5 for object detection (geetest_helpers.py:41,139). These are image-classification and object-detection models, not language models. The hcaptcha_challenger model hub is installed via solver.install(upgrade=True) (botright/botright.py:81), but this downloads vision model weights only.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-bot evasion is Botright's raison d'être, implemented across multiple layers:

Real browser. Rather than Playwright's bundled browser, it finds a locally-installed Chromium via the browsers library, preferring Ungoogled Chromium for maximum stealth (botright/botright.py:191-216).

Fingerprint spoofing. Faker.get_computer() (botright/modules/faker.py:48-52) calls chrome-fingerprints (AsyncFingerprintGenerator) to generate a realistic fingerprint. The Page._mock_page() method then overrides the user agent via CDP using Emulation.setUserAgentOverride and Network.setUserAgentOverride, setting brands, fullVersionList, platform architecture, and model (page.py:162-199).

CLI flags. ~50 Chrome flags disable detection surfaces: --disable-blink-features=AutomationControlled, --test-type, --incognito, --disable-extensions, --disable-ipc-flooding-protection, --hide-scrollbars, --safebrowsing-disable-auto-update, and many more (botright/botright.py:105-119). --ignore-default-args=["--enable-automation"] is also set (browser.py:64).

Canvas spoofing. spoof_canvas=True triggers --fingerprinting-canvas-image-data-noise for Chromium or --disable-reading-from-canvas for others, foiling canvas-fingerprinting scripts (botright/botright.py:126-131).

User-action layer. When user_action_layer=True, a transparent overlay canvas is injected showing mouse movements as colored dots — mimicking human visual feedback (page.py:201-210).

Humanized interaction. The Mouse class (mouse.py:18-248) generates Bézier-curve mouse trajectories with random distortion and tweening, with random delays between actions. The Keyboard class (keyboard.py:21-28) types characters with per-character random delays.

Proxy management. ProxyManager (proxy_manager.py:37-150) parses proxy URLs, validates them with HTTP requests against geolocation APIs, resolves the proxy's timezone/location, and passes http_credentials and browser_proxy to the browser context.

CAPTCHA solving. hCaptcha, reCaptcha, and GeeTest (v3/v4 — temporarily unavailable in released code) are solved using local computer-vision models, requiring no external services. The README reports reCaptcha v3 scores of 0.9 and CreepJS scores around 65%.

Botright does not implement rate limiting, IP rotation scheduling, or robots.txt parsing.

How is crawling at scale implemented?

not applicable

Botright does not implement any crawling-at-scale infrastructure. There are no URL queues, no URL deduplication, no crawl-depth limits, no robots.txt parsing, no politeness delays, and no distributed-worker coordination. The library provides Botright as a single-session entry point that manages one Playwright instance at a time; users who want multi-page crawling must build their own control flow on top of Botright's new_page() and page.goto() methods. The documentation explicitly notes the API is not thread-safe (docs/index.rst:208-214), recommending asyncio.gather() for concurrency.

What is the developer interface?

answered

Botright has a single interface: an asynchronous Python library. There is no CLI, REST service, MCP server, or graphical UI.

Library API — The developer imports botright and uses the async context manager pattern:

botright_client = await botright.Botright(headless=False, block_images=True, cache_responses=False, ...)
browser = await botright_client.new_browser(proxy="user:pass@ip:port")
page = await browser.new_page()
await page.goto("https://example.com")
await page.solve_hcaptcha()  # custom method
await botright_client.close()

The Botright class (botright/botright.py:27) accepts config flags (headless, block_images, cache_responses, user_action_layer, scroll_into_view, spoof_canvas, mask_fingerprint, use_undetected_playwright). new_browser() (botright/botright.py:134) takes an optional proxy string plus any Playwright launch arguments, and returns a BrowserContext (browser.py:118).

Playwright-compatible wrappers — Every standard Playwright class is wrapped with Botright's overrides: Page, BrowserContext, Locator, ElementHandle, Frame, FrameLocator, Mouse, Keyboard, Request, Response, Route (re-exported from botright/playwright_mock/__init__.py:1-12). Users can write existing Playwright code by only changing the initialization.

Extra Page methods — solve_hcaptcha(), get_hcaptcha(), solve_recaptcha(), solve_geetest() are injected directly onto the Page class (page.py:212-260).

Output formats — None built in. Content is extracted via standard Playwright methods (HTML string via page.content(), text via locator access, screenshots via page.screenshot()).

Language bindings — Python 3.8+ only. Installed via pip install botright then playwright install chromium. Dependencies include Playwright, hcaptcha_challenger, chrome-fingerprints, and recognizer (for reCaptcha).