vinyzu-archive/Botright
Archived async Playwright wrapper with real-Chrome fingerprints, humanized input and local hCaptcha/reCAPTCHA solving; no extraction.
Overview
Botright is an async Python library that wraps Playwright to make an automated browser harder to detect. You replace the browser initialisation with await botright.Botright() and await client.new_browser(proxy=...), and keep writing ordinary Playwright code. Underneath, it launches a Chromium-family browser already installed on your machine (not Playwright’s bundled build), applies a fingerprint sampled from real Chrome data, aligns timezone and geolocation with the proxy’s exit IP, moves the mouse along randomised Bézier curves, and adds page.solve_hcaptcha() and page.solve_recaptcha() methods backed by local vision models.
It is not a scraper in the extraction sense. There are no LLM calls, no HTML-to-Markdown conversion, no selectors-as-schema, and no crawler. You get the browser you would use underneath such a tool. The “AI” in its description is computer-vision CAPTCHA solving from two of the author’s other packages, hcaptcha_challenger and recognizer.
Treat the project as frozen. The pinned repository lives under a vinyzu-archive organisation, the version is 0.5.1, the last changelog entry is from January 2024, and the requirements pin playwright==1.42.0. The README marks hCaptcha solving as “outdated” and GeeTest as “Currently Not Available”. The code agrees: solve_geetest raises NotImplementedError.
Architecture
flowchart LR
U["Your async code"] --> BR["Botright client"]
BR --> ENG["get_browser_engine (local Chromium)"]
BR --> NB["new_browser(proxy)"]
NB --> PM["ProxyManager: IP + geo lookup"]
NB --> FK["Faker: chrome-fingerprints"]
NB --> CTX["launch_persistent_context"]
CTX --> BC["BrowserContext wrapper"]
BC --> PG["Page wrapper"]
PG --> CDP["CDP UA + client hints override"]
PG --> HM["Humanized Mouse / Keyboard"]
PG --> HC["hCaptcha (hcaptcha_challenger)"]
PG --> RC["reCAPTCHA (recognizer)"]
| Component | Path | Role |
|---|---|---|
| Client | botright/botright.py |
Botright async object: Chrome flags, engine discovery, new_browser, cleanup |
| Proxy manager | botright/modules/proxy_manager.py |
Parses proxy strings, checks the exit IP, looks up country, coordinates and timezone |
| Fingerprints | botright/modules/faker.py |
Gets a fingerprint from chrome-fingerprints and a locale from the proxy country |
| Context wrapper | botright/playwright_mock/browser.py |
Builds launch arguments, persistent context, image blocking, response cache |
| Page wrapper | botright/playwright_mock/page.py |
CDP UA override, CAPTCHA methods, humanised click/type overrides |
| Input | botright/playwright_mock/mouse.py, keyboard.py |
Bézier mouse paths and jittered key delays |
| hCaptcha | botright/modules/hcaptcha.py |
Drives hcaptcha_challenger.AgentT, optional rqdata injection |
| Other wrappers | botright/playwright_mock/ |
Locator, Frame, FrameLocator, handles and routes re-wrapped so behaviour carries through |
How a request flows
- Client init.
Botright.__ainit__first runshcaptcha_challenger.install(upgrade=True)to fetch or refresh model assets. It then starts Playwright (orundetected_playwrightwhenuse_undetected_playwright=True), picks a browser, deletes leftoverbotright-*temp dirs, and builds a list of about 70 Chrome flags. These include--disable-blink-features=AutomationControlledand WebRTCdisable_non_proxied_udp, plus--fingerprinting-canvas-image-data-noisewhen canvas spoofing is on (botright.py). - Engine choice.
get_browser_engineusespybrowsersto find an installed browser. It trieschromiumfirst (the author recommends Ungoogled Chromium), thenchrome, then any non-Firefox/Edge/Opera/IE browser. It raises if none is found (botright.py). - Proxy check.
new_browserawaitsProxyManager, which acceptsip:port,user:pass@ip:portand colon-separated four-part forms. It asks five what-is-my-IP services for the exit IP through the proxy, then four geo-IP services for country, latitude, longitude and timezone. With no proxy, the same lookups describe your own IP (proxy_manager.py). - Fingerprint.
Fakerconcurrently samples aChromeFingerprintand maps the proxy country code to a locale. An unknown country raisesValueError("Proxy Country not supported")(faker.py). - Launch.
browser.new_browserpasses the fingerprint’s user agent, screen and viewport, the proxy’s timezone and geolocation,color_scheme="dark", proxy credentials andignore_default_args=["--enable-automation"]tolaunch_persistent_contextwith a fresh temp profile and the local browser’s executable (browser.py). It then grants notification and geolocation permissions (botright.py). - Page setup.
new_pageopens a CDP session and sends bothEmulation.setUserAgentOverrideandNetwork.setUserAgentOverridewith full client-hints metadata (brands, full version list, platform, architecture, bitness, model) from the fingerprint (page.py). - Interaction.
page.click(selector)waits for the element, takes its bounding box, and callsMouse.click. That moves along a humanised curve from the last cursor position, pauses 200 to 400 ms, then presses and releases (page.py, mouse.py).
Key components
Playwright mock layer
Each wrapper subclasses the Playwright class, adopts the original’s _impl_obj, keeps the original methods as _origin_*, and overrides the ones that return other Playwright objects. That way objects reached through page.locator(...) or page.frames are Botright types too. Route handlers are proxied so callbacks receive wrapped Route/Request objects. With use_undetected_playwright, methods that depend on the CDP Runtime domain (expose_function, expose_binding, expect_console_message) raise NotSupportedError (browser.py). This works, but it is coupled tightly to Playwright internals, which is one reason the dependency is pinned.
Human-like input
HumanizeMouseTrajectory (adapted from HumanCursor) picks two random knots within 80 px of the straight path, builds a Bézier curve, adds Gaussian distortion, and tweens to 100 points with an ease-out. Each point becomes one mouse.move (mouse.py). Keyboard.type sends one character at a time with 50 to 150 ms delays by default (keyboard.py). Humanisation is always on. The user_action_layer flag only injects a full-screen canvas that draws coloured dots where the mouse moves and clicks. It is a debugging overlay, not a stealth feature.
CAPTCHA solving
hCaptcha.solve_hcaptcha clicks the checkbox through AgentT, retries up to seven rounds (refreshing on a backcall), and returns the pass token. If you give it rq_data, it first intercepts hcaptcha.com/getcaptcha/** and re-posts the request with that rqdata (hcaptcha.py). On failure it returns an error string (“Exceeded maximum retry times…”), not None, so check the result before using it as a token. solve_recaptcha delegates to recognizer’s AsyncChallenger. solve_geetest raises NotImplementedError, even though a 368-line geetest.py and a TorchScript model still ship in the package (page.py).
Bandwidth helpers
block_images=True aborts image requests by extension and by resource type. cache_responses=True serves documents, stylesheets, images, fonts and media from an in-memory dict keyed by URL, shared across all browsers of one client (browser.py). Caching document responses means a second goto to the same URL replays stale HTML.
Extending it
- Existing Playwright code. Everything after
new_browser()is the Playwright async API, so extraction is up to you:page.content(), locators, or a downstream LLM extractor. - Launch overrides. Extra keyword arguments to
new_browserare merged over Botright’s launch arguments, so you can overridelocale,viewportorcolor_scheme. - Stealth knobs.
mask_fingerprint=Falsekeeps the real UA and screen (the README says one tough site only passes this way).spoof_canvas,headlessanduse_undetected_playwrightare constructor flags. - Concurrency. Not thread-safe. Use one client per thread, or several pages and browsers with
asyncio.gather.
Running it
pip install botright, thenplaywright install. You also need a locally installed Chromium or Chrome; Ungoogled Chromium is recommended for canvas noise.- Python 3.8+, async only. The default is a visible window (
headless=False). - The first run downloads
hcaptcha_challengerassets, and every client start callsinstall(upgrade=True)again. - Each
new_browsercall makes outbound requests to public IP and geo-IP services (ipify, ip-api.com, ipapi.co and others). Plan for those in locked-down networks.
Strengths and caveats
- Strength: coherent identity. UA string, client hints, screen, timezone and geolocation all come from one fingerprint and the proxy’s real exit location, rather than being randomised independently.
- Strength: drop-in API. Thorough wrapping means existing Playwright scripts mostly just work, with humanised input for free.
- Strength: local CAPTCHA solving. hCaptcha and reCAPTCHA image challenges are attempted with no paid solver API.
- Caveat: abandoned and pinned. The repository is archived, it is tied to Playwright 1.42 internals, GeeTest is disabled, and hCaptcha is flagged as outdated. Detection vendors have moved on since 2024, so the README’s pass table is historical.
- Caveat: locale inconsistency. The browser locale and
Accept-Languageare hard-coded toen-US, whatever the proxy country. The computed per-country locale is never applied, and many of its codes are invalid anyway (sp-ES,po-BR,ar-IL). A German proxy therefore shows Berlin time with US English. - Caveat: small bugs.
Locator.clickreads the bounding box beforescroll_into_view_if_neededand clicks the stale coordinates.solve_hcaptchareturns an error string on failure. An unsupported proxy country abortsnew_browser. - Caveat: not an extraction tool. No LLM, content cleaning, crawling, rate limiting or proxy rotation. Pair it with something else for the data side.
Sources: code at 6e64fed, deepwiki-open wiki (12 pages), OpenDeepWiki wiki (9 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 6e64fed. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredBotright uses headless Chromium via Playwright exclusively — there is no plain-HTTP fetch path. The Botright.__ainit__ method (botright/botright.py:55) starts Playwright, then get_browser_engine() (botright/botright.py:191-216) locates a locally-installed Chromium-based browser (preferring Ungoogled Chromium, falling back to Chrome, then any Chromium variant). Browsers are launched as persistent contexts via launch_persistent_context() (botright/playwright_mock/browser.py:88-95), which preserves session state across pages.
JS rendering is handled automatically by Playwright's browser engine — there is no toggle. Pages are created through the Page wrapper (botright/playwright_mock/page.py:58-74), which calls _mock_page() to apply stealth patches via CDP (page.py:162-210), then returns the page.
The waiting strategy delegates to standard Playwright waits — wait_for_selector, wait_for_timeout. Custom interaction methods (page.click, page.type) add humanized mouse movement with Bézier curves and random delays, but these are for stealth rather than content-waiting.
Content types: PDF and image handling is left entirely to Playwright (e.g. page.screenshot()). Image blocking is available via the block_images flag (botright/botright.py:36), which installs a route interceptor aborting image resource types (browser.py:191-202). Response caching for static resources (document, stylesheet, image, media, font, manifest) is supported through cache_responses (browser.py:175-189), storing APIResponse objects in a dict keyed by URL.
There is no specialized PDF rendering or waiting beyond Playwright's built-in capabilities.
headless=False), launched from a locally installed Chromium/Chrome executable found via pybrowsers, not Playwright's bundled headless Chromium (botright/botright.py L55-L72, L191-L216).How is content extracted or converted?
not applicableBotright does not implement any content-extraction or conversion utilities. There is no HTML→Markdown conversion, no readability-style boilerplate removal, no CSS selector-based extraction schema, and no structured-output pipeline. The library exposes the full Playwright Page API (page.content(), page.inner_text(), page.text_content(), etc.) for users to extract whatever they need, but provides no helpers beyond that. Its purpose is undetected browser automation and CAPTCHA solving, not content parsing or transformation.
How are LLMs used, if at all?
not applicableBotright does not use LLMs anywhere in its codebase. There are no prompts, no page-chunking strategies, no structured output schemas for LLM consumption, no cost controls, and no integration with any LLM provider (OpenAI, Anthropic, or otherwise). The 'AI' components mentioned are entirely computer-vision models used for CAPTCHA solving: hcaptcha_challenger uses vision models for hCaptcha image classification; recognizer/reCognizer handles reCaptcha; the commented-out GeeTest solver used CLIP-ViT-B-32 (a vision transformer) via SentenceTransformer for image similarity and YOLOv5 for object detection (geetest_helpers.py:41,139). These are image-classification and object-detection models, not language models. The hcaptcha_challenger model hub is installed via solver.install(upgrade=True) (botright/botright.py:81), but this downloads vision model weights only.
How are anti-bot measures, proxies and fingerprinting handled?
answeredAnti-bot evasion is Botright's raison d'être, implemented across multiple layers:
Real browser. Rather than Playwright's bundled browser, it finds a locally-installed Chromium via the browsers library, preferring Ungoogled Chromium for maximum stealth (botright/botright.py:191-216).
Fingerprint spoofing. Faker.get_computer() (botright/modules/faker.py:48-52) calls chrome-fingerprints (AsyncFingerprintGenerator) to generate a realistic fingerprint. The Page._mock_page() method then overrides the user agent via CDP using Emulation.setUserAgentOverride and Network.setUserAgentOverride, setting brands, fullVersionList, platform architecture, and model (page.py:162-199).
CLI flags. ~50 Chrome flags disable detection surfaces: --disable-blink-features=AutomationControlled, --test-type, --incognito, --disable-extensions, --disable-ipc-flooding-protection, --hide-scrollbars, --safebrowsing-disable-auto-update, and many more (botright/botright.py:105-119). --ignore-default-args=["--enable-automation"] is also set (browser.py:64).
Canvas spoofing. spoof_canvas=True triggers --fingerprinting-canvas-image-data-noise for Chromium or --disable-reading-from-canvas for others, foiling canvas-fingerprinting scripts (botright/botright.py:126-131).
User-action layer. When user_action_layer=True, a transparent overlay canvas is injected showing mouse movements as colored dots — mimicking human visual feedback (page.py:201-210).
Humanized interaction. The Mouse class (mouse.py:18-248) generates Bézier-curve mouse trajectories with random distortion and tweening, with random delays between actions. The Keyboard class (keyboard.py:21-28) types characters with per-character random delays.
Proxy management. ProxyManager (proxy_manager.py:37-150) parses proxy URLs, validates them with HTTP requests against geolocation APIs, resolves the proxy's timezone/location, and passes http_credentials and browser_proxy to the browser context.
CAPTCHA solving. hCaptcha, reCaptcha, and GeeTest (v3/v4 — temporarily unavailable in released code) are solved using local computer-vision models, requiring no external services. The README reports reCaptcha v3 scores of 0.9 and CreepJS scores around 65%.
Botright does not implement rate limiting, IP rotation scheduling, or robots.txt parsing.
How is crawling at scale implemented?
not applicableBotright does not implement any crawling-at-scale infrastructure. There are no URL queues, no URL deduplication, no crawl-depth limits, no robots.txt parsing, no politeness delays, and no distributed-worker coordination. The library provides Botright as a single-session entry point that manages one Playwright instance at a time; users who want multi-page crawling must build their own control flow on top of Botright's new_page() and page.goto() methods. The documentation explicitly notes the API is not thread-safe (docs/index.rst:208-214), recommending asyncio.gather() for concurrency.
What is the developer interface?
answeredBotright has a single interface: an asynchronous Python library. There is no CLI, REST service, MCP server, or graphical UI.
Library API — The developer imports botright and uses the async context manager pattern:
botright_client = await botright.Botright(headless=False, block_images=True, cache_responses=False, ...)
browser = await botright_client.new_browser(proxy="user:pass@ip:port")
page = await browser.new_page()
await page.goto("https://example.com")
await page.solve_hcaptcha() # custom method
await botright_client.close()
The Botright class (botright/botright.py:27) accepts config flags (headless, block_images, cache_responses, user_action_layer, scroll_into_view, spoof_canvas, mask_fingerprint, use_undetected_playwright). new_browser() (botright/botright.py:134) takes an optional proxy string plus any Playwright launch arguments, and returns a BrowserContext (browser.py:118).
Playwright-compatible wrappers — Every standard Playwright class is wrapped with Botright's overrides: Page, BrowserContext, Locator, ElementHandle, Frame, FrameLocator, Mouse, Keyboard, Request, Response, Route (re-exported from botright/playwright_mock/__init__.py:1-12). Users can write existing Playwright code by only changing the initialization.
Extra Page methods — solve_hcaptcha(), get_hcaptcha(), solve_recaptcha(), solve_geetest() are injected directly onto the Page class (page.py:212-260).
Output formats — None built in. Content is extracted via standard Playwright methods (HTML string via page.content(), text via locator access, screenshots via page.screenshot()).
Language bindings — Python 3.8+ only. Installed via pip install botright then playwright install chromium. Dependencies include Playwright, hcaptcha_challenger, chrome-fingerprints, and recognizer (for reCaptcha).