LLMs Technical Reviews
Home / AI web scraping / Selenium-Driverless

ttlns/Selenium-Driverless

Async Python driver that runs real Chrome over CDP without chromedriver, built to avoid bot detection; no AI or extraction.

GitHub ↗★ 861PythonCC-BY-NC-SA-4.0 (commercial licence sold separately)commit ca333fd · 2024-11-23homepage ↗

Overview

Selenium-Driverless is not an AI scraper. It is a Python browser driver. It starts a normal Chrome binary with --remote-debugging-port and controls it over the Chrome DevTools Protocol (CDP) from an asyncio client. It does not use chromedriver or the W3C WebDriver protocol. The API copies Selenium’s names (driver.get, find_element(By.CSS_SELECTOR, ...), execute_script), but nothing in the package talks to Selenium’s wire protocol.

The project exists to avoid bot detection. Without chromedriver, there are no chromedriver page artifacts and no --enable-automation switch. Your scripts run in a CDP isolated world, so page JavaScript cannot see them. Clicks come from Input.dispatchMouseEvent with curved, timed mouse paths. The README claims it passes Cloudflare, Bet365 and Turnstile. The tests/antibots/ folder has tests against those targets and a few detector pages. We did not run them.

Where it fits in an AI scraping stack: it is the fetch layer. You use it to get the rendered HTML of pages that block Playwright or plain HTTP clients. The package has no LLM code, no HTML-to-text conversion, no crawl queue and no URL dedup. page_source returns raw HTML. You pass that HTML to a converter (for example Turndown or a readability library) and then to your model. A crawler or job platform drives the URLs.

The licence matters. LICENSE.md is CC BY-NC-SA 4.0 plus a custom clause. Commercial use is reserved and needs a paid licence from the author. The first run prints the licence to stderr.

Architecture

flowchart LR
  U["Your asyncio code"] --> C["webdriver.Chrome"]
  C -->|"subprocess.Popen"| CH["Chrome binary"]
  C --> BT["BaseTarget (browser CDP socket)"]
  C --> CTX["Context (default / incognito)"]
  CTX --> T["Target (one tab or iframe)"]
  T --> WE["WebElement / JSRemoteObj"]
  T --> P["Pointer (human-like mouse)"]
  T --> NI["NetworkInterceptor (Fetch domain)"]
  C --> EXT["MV3 helper extension"]
  EXT -->|"chrome.proxy / webRequest"| CH
  BT --> CH
  T --> CH
Component Path Role
Driver src/selenium_driverless/webdriver.py Chrome class: launches Chrome, discovers targets, manages contexts, proxy and auth
Options src/selenium_driverless/types/options.py Default Chrome flags, profile dir, headless, proxy, extensions
Target src/selenium_driverless/types/target.py One CDP page or iframe target: get, scripts, cookies, fetch/xhr, screenshots
Context src/selenium_driverless/types/context.py A browser context (default or incognito) with its own tabs and storage
Remote objects src/selenium_driverless/types/deserialize.py Wraps JS objects. Creates the isolated world used for scripts
Elements src/selenium_driverless/types/webelement.py Locators, click, send_keys, box model
Input src/selenium_driverless/input/pointer.py, scripts/geometry.py Mouse paths and click timing
Interception src/selenium_driverless/scripts/network_interceptor.py Request/response/auth interception on the CDP Fetch domain
Helper extension src/selenium_driverless/files/mv3_extension/ MV3 service worker that gives the driver chrome.proxy, webRequest and privacy access
Sync wrapper src/selenium_driverless/sync/ Blocking API that runs each coroutine on a private event loop

How a request flows

Take async with webdriver.Chrome() as driver: await driver.get(url):

  1. Launch. Chrome.start_session adds the helper extension and picks a random free port. It writes Preferences into a temporary --user-data-dir and starts Chrome with subprocess.Popen. In headless mode it adds --user-agent= with the UA saved from an earlier headful run. Without a saved UA it warns that “headless is detectable at first run” (webdriver.py).
  2. Attach. It opens a BaseTarget socket to the browser endpoint. It saves Browser.getVersion’s UA with HeadlessChrome changed to Chrome. It attaches to the first non-extension page target and wraps it in a Context. Then it turns on focus emulation and applies single_proxy and download behaviour (webdriver.py).
  3. Navigate. Target.get enables the Page domain. It races Page.loadEventFired against a download event, then sends Page.navigate with transitionType: "link" and an optional referrer. Fragment-only URLs skip the wait (target.py). There is no network-idle wait. You wait for elements yourself (find_element(..., timeout=...)).
  4. Script and locate. execute_script runs by default in an isolated world. Page.createIsolatedWorld creates it with universal access (deserialize.py, target.py). Locators run querySelectorAll or document.evaluate in that world. By.ID, By.NAME and By.CLASS_NAME are first rewritten to XPath (webelement.py).
  5. Interact. WebElement.click scrolls into view and picks a random point inside the box model (webelement.py). Pointer.move_to builds a curved path with gen_combined_path and plays it back at about 60 Hz. The press-to-release gap is about 125 ms with jitter (pointer.py, L182-L195).
  6. Read. page_source returns the document’s outer HTML. It retries on stale references for up to 10 s (target.py). This HTML is what you pass to the next stage.

Key components

Launch flags and the user agent

Options.__init__ adds a set of quiet flags: --no-first-run, background-throttling switches, --disable-background-networking, and --enable-privacy-sandbox-ads-apis to restore APIs that Chrome hides under automation (options.py). Setting headless adds --headless=new. The library has no fingerprint-spoofing layer: no canvas, WebGL or font patches. The approach is “use real Chrome and avoid leaving traces”. The README points to the separate CDP-Patches project for input-level leaks.

Contexts and proxies

new_context calls Target.createBrowserContext with an optional proxyServer and bypass list. Each incognito context has its own cookies and storage (webdriver.py). Global proxies use the helper extension instead. set_single_proxy parses scheme://user:pass@host:port and calls chrome.proxy.settings.set inside the extension’s service worker. Credentials go into a webRequest.onAuthRequired handler (webdriver.py). Proxy changes take effect at runtime, without a browser restart. The same extension target also sets webRTCIPHandlingPolicy to disable_non_proxied_udp, so WebRTC does not leak the real IP (webdriver.py). There is no proxy pool or rotation. You call set_single_proxy or open a new context per proxy.

Helper extension

The extension’s background script is only a keep-alive loop that pings a port every 25 s (background js). Its value is its privileged context. The driver attaches to the service worker over CDP and runs chrome.* calls there. It does not collect fingerprints.

Network interception and in-page requests

NetworkInterceptor is an async context manager over CDP Fetch.enable. It sends request, response and auth events to callbacks or to an async for iterator, and resumes any request the callback did not handle (network_interceptor.py). Target.fetch and Target.xhr run a JS fetch/XHR inside the page. The request carries the page’s cookies, TLS fingerprint and origin, which is useful for calling a site’s JSON API after a challenge (target.py).

Sync wrapper

sync.webdriver.Chrome overrides __getattribute__. Coroutine methods run with loop.run_until_complete on a private loop (sync/webdriver.py). It is convenient, but you cannot use it inside a running event loop.

Extending it

  • Raw CDP. execute_cdp_cmd, add_cdp_listener and get_cdp_event_iter expose any CDP domain on the driver, a context or a target.
  • Interception hooks. Callbacks get InterceptedRequest objects with continue_request, fulfill, fail and body access. You can block assets, rewrite headers or capture API responses.
  • Your own extensions. options.add_extension(path_or_zip) loads more unpacked or zipped extensions next to the helper.
  • Remote browsers. Set options.debugger_address to attach to a Chrome you started yourself.

Running it

  • pip install selenium-driverless. Python 3.8 or newer, with cdp-socket, numpy, scipy, aiofiles and platformdirs from setup.py. Google Chrome must be installed. find_chrome_executable locates it, or you set options.binary_location.
  • There is no server or service. Each Chrome() is one local browser process with a temp profile, removed on quit() unless auto_clean_dirs is off.
  • The project’s own tests run headful (no_headless = True in tests/conftest.py). Headless works but is the weaker mode for stealth. On Linux servers, use Xvfb or a similar virtual display.

Strengths and caveats

  • Strength: real Chrome, no driver artifacts. It uses CDP only and runs scripts in isolated worlds. This avoids the most common automation tells without patching the browser.
  • Strength: proxy features. Per-context proxies, runtime proxy switching with auth, and the WebRTC policy are hard to set up by hand in other tools.
  • Strength: in-page fetch and Fetch interception. You can reuse a cleared session for API calls and capture JSON instead of parsing HTML.
  • Caveat: not AI, not a crawler. No extraction, Markdown output, queue, concurrency control or robots.txt handling. Treat it as a replacement for Playwright’s page.goto and build the rest yourself.
  • Caveat: licence. CC BY-NC-SA with a paid commercial licence. This rules it out for many company stacks.
  • Caveat: rough edges in the code. execute_raw_script checks while elapsed > timeout, which is backwards. The loop never runs, and the method always raises TimeoutError (target.py). A missing comma in the default flags joins --homepage=about:blank and --wm-window-animations-disabled into one flag (options.py). By.CLASS_NAME matches the whole class attribute exactly, so an element with several classes is not found by one of them.
  • Caveat: momentum. The pinned commit is from November 2024, and setup.py still says “Pre-Alpha”. Stealth tools age quickly as detectors change. Test against your targets before you rely on it.

Sources: code at ca333fd, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (16 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit ca333fd. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Pages are fetched and rendered exclusively via a full headless (or headed) Chrome browser controlled through the Chrome DevTools Protocol over WebSocket — there is no plain-HTTP fetch mode. The core workflow in Target.get() (src/selenium_driverless/types/target.py:319-377) sends a Page.navigate CDP command and waits for Page.loadEventFired (or a download event) to confirm the page finished loading; setting wait_load=True configures an asyncio.wait across both events. JS rendering is native because Chrome itself executes JavaScript — the project is a driver, not a fetcher. The webdriver.Chrome class (src/selenium_driverless/webdriver.py:63-65) manages the Chrome process (launched as a subprocess with --remote-debugging-port), connects via WebSocket to CDP, and surfaces Target objects per tab. Additionally, Target.fetch() (src/selenium_driverless/types/target.py:1130-1245) runs a JavaScript fetch call inside the browser (not a Python HTTP request), returning the binary body and headers as base64-decoded Python dicts. Target.xhr() does the same via XMLHttpRequest. For PDFs and images, the browser's native download capability is used: wait_download() (src/selenium_driverless/types/target.py:272-316) listens for Browser.downloadWillBegin CDP events and resolves to the local file path when downloads_dir is configured. print_page() returns PDF via Page.printToPDF, and get_screenshot_as_png() returns screenshots via Page.captureScreenshot. The NetworkInterceptor (src/selenium_driverless/scripts/network_interceptor.py:629-800) can intercept requests/responses at the CDP Fetch level, allowing body inspection, header modification, or external fetch bypass.

How is content extracted or converted?

answered

Content extraction is rudimentary — there is no built-in HTML→Markdown conversion, readability-style boilerplate removal, or schema-based extraction. The library exposes raw DOM access and expects the consuming code to handle further processing. Target.page_source (src/selenium_driverless/types/target.py:649-661) retrieves the full outer HTML of the document via DOM.getOuterHTML. WebElement.source (src/selenium_driverless/types/webelement.py:373-385) does the same per element. WebElement.text (line 419-421) returns elem.textContent (the JavaScript property). WebElement.value (line 424-426) returns elem.value. The By selector class (src/selenium_driverless/types/by.py:22-37) supports ID, NAME, XPATH, TAG_NAME, CLASS_NAME, and CSS_SELECTOR. Elements are found through DOM node IDs (DOM.getDocument via _document_elem at line 840-852, then find_element recursively walking the DOM tree). Target.search_elements() (src/selenium_driverless/types/target.py:913-942) uses CDP's DOM.performSearch for DevTools-style find-in-page. For shadow DOM: WebElement.shadow_root() (line 264) and WebElement.content_document() (line 199) provide access to shadow roots and iframe content documents, letting scrapers navigate complex widget trees — demonstrated in the Turnstile bypass test (tests/antibots/test_cloudfare.py). Snapshots are available as MHTML via Target.snapshot(). There is no readability, no structured output schema, no markdown conversion, and no content cleaning pipeline.

How are LLMs used, if at all?

not applicable

The repository contains no LLM integration whatsoever. There are no imports, references, or integrations with OpenAI, Anthropic, LangChain, or any LLM provider anywhere in the source code (src/). A string search for openai, anthropic, gpt, langchain, llm, markdown, readability, and boilerplate across all Python source files returns only a false positive (re.fullmatch matching on "llm" substring). There is no prompt management, no chunking logic for large pages, no structured-output parser, and no cost-control mechanism. The project is a pure browser-automation driver, not an AI-extraction pipeline. If a user wants to extract structured data, they must supply their own LLM or scraping logic on top.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-bot bypass is the project's headline feature. The fundamental approach is eliminating chromedriver — the process that anti-bot scripts typically detect. Instead, the library talks CDP directly via WebSocket (cdp-socket). Key techniques: (1) User-agent patching — on first headless run (src/selenium_driverless/webdriver.py:163-168) the detected UA string has "HeadlessChrome" stripped to "Chrome" and is cached to disk for reuse. (2) Human-like mouse movements — geometry.py generates Bezier-spline mouse paths with cubic easing and Gaussian noise via NumPy/SciPy; Pointer.move_to() uses gen_combined_path with configurable smoothness and acceleration. (3) Focus emulation — Emulation.setFocusEmulationEnabled(true) is called on session start (line 305) to prevent detection of unfocused tabs. (4) CDP patches — an external cdp-patches package (referenced in tests/antibots/test_brotector.py:2) integrates additional headful-only stealth patches. (5) MV3 Chrome extension (src/selenium_driverless/files/mv3_extension/) — loaded as a service_worker with broad permissions including proxy, webRequest, webRequestBlocking, and cookies. It handles proxy auth (set_auth), dynamic proxy switching (set_single_proxy, clear_proxy), and WebRTC IP-leak prevention. (6) Dynamic proxy — options.single_proxy and driver.set_single_proxy() configure per-context Chrome proxy settings via the extension's chrome.proxy.settings.set API (webdriver.py:1317-1378). (7) Per-context isolation — incognito contexts with isolated cookies and local storage via Target.createBrowserContext. (8) CAPTCHA handling — not automated, but the Turnstile bypass test (tests/antibots/test_cloudfare.py) demonstrates navigating Cloudflare Turnstile through iframe/shadow-DOM traversal with mouse movements, then clicking the checkbox. There is no rate-limiting, no proxy rotation, and no automatic CAPTCHA-solver integration — those are left to the user.

How is crawling at scale implemented?

not applicable

This project does not implement crawling at scale. There are no URL queues, no deduplication, no depth/limit controls, no robots.txt parser, no politeness delays, no distributed workers, and no rate-limiting — none of these terms appear in the source code. The library is designed as a single-session browser driver for one-off page interactions or small-scale scraping tasks where anti-bot stealth is paramount. Crawling logic (managing multiple URLs, concurrent contexts, scheduling) must be built externally by the user.

What is the developer interface?

answered

The developer interface is exclusively a Python library — there is no CLI, no REST API, no MCP server, and no graphical UI. Two entry points serve usage preferences. Async API: from selenium_driverless import webdriver — a conventional asyncio-based async with pattern (src/selenium_driverless/webdriver.py:63-65). Sync API: from selenium_driverless.sync import webdriver (src/selenium_driverless/sync/webdriver.py:8-40) — the synchronous version runs an internal event loop and wraps every coroutine via loop.run_until_complete, so methods like driver.get() and driver.find_element() can be called without await. The API mirrors Selenium's naming conventions: chrome.get(), chrome.find_element(By, value), chrome.page_source, chrome.current_url, etc. Context management — the library adds new_context() (incognito contexts with isolated cookies/storage) and switch_to for multi-tab workflows. Network interception — the NetworkInterceptor context manager (src/selenium_driverless/scripts/network_interceptor.py:629-800) provides on_request, on_response, and on_auth callbacks with methods to continue_request, fulfill, fail_request, and bypass_browser (re-fetch from Python). It also supports async for iteration over intercepted requests. Output formats — content is returned as Python dicts or WebElement objects. No JSON Lines, no file dumps, no streaming output. Page source is raw HTML (string). Screenshots are PNG (bytes). PDFs are base64 strings. Snapshots are MHTML (string). Language bindings — Python only.