LLMs Technical Reviews
Home / AI web scraping / spider-flow

ssssssss-team/spider-flow

Java/Spring Boot visual flowchart crawler builder on Jsoup and selector expressions; no JS rendering or LLMs, last commit 2022.

GitHub ↗★ 11kJavaMITcommit c799cca · 2022-10-25homepage ↗

Overview

Spider-Flow is a visual crawler builder, not an AI scraper. It is a Spring Boot 2.0 web app (Java 8). You draw a scraping job as a flowchart in the browser: a start node, request nodes, variable and loop nodes, SQL and output nodes. Arrows between nodes carry conditions. Each node holds small expressions such as ${extract.xpath(resp.element(), '//h1')}. The server runs the chart on a thread pool, fetches pages with Jsoup, and writes rows to a database or CSV. It has no LLM code and no headless browser.

The project’s last commit on the default branch (the one reviewed here) is from October 2022, version 0.5.0. Dependencies are from that period: Spring Boot 2.0.7, Nashorn for user-defined JavaScript functions (removed from the JDK in Java 15), and FastJson. Selenium rendering, proxy pools, OCR, Redis and Mongo outputs are separate plugins in other Gitee repositories. They are not in this codebase.

Where it fits in an AI scraping stack: it is a low-code fetch-and-select tool for simple, server-rendered sites, used by people who prefer a flowchart to code. It is a weak fit for modern AI pipelines. It cannot render JavaScript without the external plugin. It has no HTML-to-Markdown step, so you cannot easily send clean text to a model. The UI has no authentication. If you only need a scheduler around scripts, Crawlab is the closer fit. If you need LLM-ready output, use a newer crawler.

Architecture

flowchart LR
  ED["Browser editor (mxGraph)"] -->|"flow XML"| WEB["Spring MVC controllers"]
  ED -->|"WebSocket /ws test/debug"| WS["WebSocketEditorServer"]
  WEB --> DB["MySQL (flows, tasks, data sources)"]
  CRON["Quartz SpiderJob"] --> SP["Spider engine"]
  WEB -->|"/rest/run"| SP
  WS --> SP
  SP --> POOL["Thread pool + submit strategy"]
  POOL --> EX["Shape executors"]
  EX --> REQ["RequestExecutor (Jsoup)"]
  EX --> OUT["Output / SQL executors"]
  EX --> EXPR["Expression engine + functions"]
  OUT --> RES["JDBC tables / CSV"]
Component Path Role
API module spider-flow-api/ Interfaces: ShapeExecutor, FunctionExecutor, FunctionExtension, PluginConfig, SpiderContext
Engine spider-flow-core/.../core/Spider.java Parses flow XML, walks nodes, runs loops, manages futures
Thread pool spider-flow-api/.../concurrent/ Global pool with per-flow sub-pools and four submit strategies
Shape executors spider-flow-core/.../executor/shape/ RequestExecutor, VariableExecutor, LoopExecutor, ForkJoinExecutor, OutputExecutor, ExecuteSQLExecutor, ProcessExecutor, and more
Functions spider-flow-core/.../executor/function/ extract, string, date, json, file, md5 and other expression prefixes, plus type extensions (resp.xpath, element.selector)
Expressions spider-flow-core/.../expression/ Template engine with its own tokenizer, parser and reflection-based interpreter
HTTP spider-flow-core/.../io/HttpRequest.java Thin wrapper over a Jsoup Connection
Scheduling spider-flow-core/.../job/ Quartz SpiderJob and SpiderJobManager
Web spider-flow-web/ Controllers, WebSocket debugger, static LayUI + mxGraph UI, application.properties

How a request flows

Take a flow “start → request list page → loop over links → request detail → output”:

  1. Trigger. A cron fire (SpiderJob.executeInternal), a REST call (/rest/run/{id} or /rest/runAsync/{id}) or the editor’s Test button (WebSocket test/debug event) calls Spider.run or runWithTest (SpiderJob.java, SpiderRestController.java, WebSocketEditorServer.java).
  2. Load and pool. run parses the stored mxGraph XML into a SpiderNode tree. executeRoot reads the start node’s thread count (default 8) and submit strategy (random, linked, child-first or parent-first). It then creates a sub-pool of the global 64-thread executor (Spider.java).
  3. Schedule loop. One coordinator task runs the root, then polls the future queue every millisecond. When a future is done (picked by the strategy’s comparator), it calls allowExecuteNext and then executeNextNodes on the children (Spider.java).
  4. Execute a node. executeNode checks the arrow’s condition and exception-flow flag. It evaluates the node’s loop count or collection and creates one task per iteration. Each task gets a copy of the variables, with the index and item added. Async shapes go to the pool, and sync shapes run inline. A thrown exception is stored as ex, so later arrows can branch on failure (Spider.java).
  5. Request. RequestExecutor.execute applies the per-node sleep and checks the optional Bloom filter. It builds the URL, method, headers, cookies, body and host:port proxy from expressions, then calls HttpRequest.execute. That call uses Jsoup with ignoreContentType, ignoreHttpErrors and an unlimited body size. Only HTTP 200 counts as success. Anything else is retried up to the configured count. On success the response is stored as the resp variable (RequestExecutor.java, HttpRequest.java).
  6. Extract. Variable and output nodes evaluate expressions such as ${extract.selector(resp.html,'div > a','attr','href')} or ${resp.jsonpath('$.data')}. These call Jsoup CSS selectors, Xsoup XPath, regex or FastJson JSONPath (ExtractFunctionExecutor.java).
  7. Output. OutputExecutor evaluates each name/value pair. It adds the pairs to the run’s outputs (returned as JSON by /rest/run) and can insert a row into a configured JDBC data source or append to a CSV (OutputExecutor.java).

Key components

The flow engine

Flows are graphs, not strict DAGs. A node can loop back, so test runs count executions and stop past spider.detect.dead-cycle (default 5000) (Spider.java). Production runs have no such guard. The coordinator’s 1 ms polling loop scans the whole future queue on each pass. That is fine for hundreds of pending requests and wasteful for many thousands. A ForkJoin shape uses per-node counters to wait for all branches. ProcessExecutor runs another flow as a sub-routine.

Request node

All fetching is plain HTTP through Jsoup. Each node has timeout (default 60 s), method, redirect and TLS-validation toggles, form or raw bodies, file uploads, a response charset override, and retry count and interval. Cookies set by responses go into a shared CookieContext and are re-sent by default. Politeness is manual: a per-node sleep expression enforces a minimum gap between that node’s requests. There is no robots.txt handling, user-agent rotation or proxy pool. Proxy strings without credentials are the only proxy option.

URL dedup

Dedup is opt-in per request node. When it is on, createBloomFilter loads <workspace>/<flowId>/url.bf or creates a Guava Bloom filter (capacity 1,000,000, error rate 0.0001 by default). afterEnd writes it back to disk, so dedup carries across runs (RequestExecutor.java). A URL is recorded only after a 200 response, so failed URLs are retried on the next run.

Expression language

DefaultExpressionEngine builds a context with every FunctionExecutor under its prefix and every global variable. It then renders the ${...} template (DefaultExpressionEngine.java). The interpreter calls methods by reflection. FunctionExtension classes add methods to existing types, such as resp.xpath(...) or list.join(...). User-defined functions are JavaScript, compiled by Nashorn through ScriptManager.

Extending it

  • New node types. Implement ShapeExecutor (supportShape, execute, optional allowExecuteNext/isThread) as a Spring bean. ExecutorsUtils finds it by shape name (ShapeExecutor.java).
  • New functions. Implement FunctionExecutor with a prefix and public static methods. @Comment and @Example feed the editor’s autocomplete.
  • Plugins. The external Selenium, proxy-pool, OCR, Redis, OSS, Mongo and mailbox plugins are jars that add executors and register a PluginConfig.
  • Listeners. SpiderListener.beforeStart/afterEnd hooks run around each flow. The Bloom filter persistence uses them.

Running it

  • Build with Maven (spider-flow-web produces spider-flow.jar). Import db/spiderflow.sql into MySQL and set the datasource in application.properties. The shipped values are root / 123456789 on localhost:3306. The app listens on 8088. The Dockerfile uses java:8.
  • Scheduled flows do nothing until you set spider.job.enable=true. The shipped properties file sets it to false, and SpiderJob.executeInternal returns early (SpiderJob.java).
  • Use a Java 8–14 runtime. Custom JavaScript functions need Nashorn.
  • It is a single process. There are no distributed workers, and all concurrency is threads inside one JVM.

Strengths and caveats

  • Strength: approachable. A non-programmer can build a paginated list → detail → database job in the editor, test it step by step over WebSocket and schedule it.
  • Strength: useful primitives. XPath, CSS, regex and JSONPath in one expression language, plus loop/fork-join nodes, cookie carry-over and persistent Bloom-filter dedup.
  • Caveat: not AI, no rendering. It has no LLM calls, no Markdown conversion and no JS execution without the external Selenium plugin. Its output is selector-defined fields only.
  • Caveat: no authentication. There is no login and no security filter. Anyone who reaches the port can run flows, use the WebSocket test endpoint, or run SQL against configured data sources through ExecuteSQLExecutor. Keep it on a private network.
  • Caveat: unmaintained stack. Spring Boot 2.0, Java 8, Nashorn and an old FastJson make upgrades and security patching your job.
  • Caveat: brittle success rule. Only status 200 counts as success. A 201, 204 or unfollowed 3xx is retried and then logged as a failure.

Sources: code at c799cca, deepwiki-open wiki (13 pages), OpenDeepWiki wiki (12 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit c799cca. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

All HTTP requests go through the HttpRequest class (spider-flow-core/.../io/HttpRequest.java), which wraps Jsoup's Connection - a pure-Java HTTP client, not a headless browser. The RequestExecutor shape (spider-flow-core/.../RequestExecutor.java) configures each request: method (GET/POST etc.), URL (via expression evaluation), timeout (default 60s, "timeout" field), follow-redirect toggle, TLS-certificate validation toggle, custom headers, and per-node/project-level cookies. JS rendering is not built in; the README links a separate Selenium plugin (spider-flow-selenium). The Jsoup connection calls ignoreContentType(true) and ignoreHttpErrors(true) with maxBodySize(0) (unlimited) in HttpRequest.execute() (lines 121-128). The response is wrapped as HttpResponse (spider-flow-core/.../io/HttpResponse.java) implementing the SpiderResponse interface. It exposes getHtml() (raw body), getBytes() (byte array), getStream() (InputStream), getJson() (parsed via FastJson), getContentType(), and getStatusCode() - so all content types (HTML, JSON, binary) are captured. A setCharset() override is available per node (RequestExecutor line 267-271). The waiting/politeness strategy is manual: each request node can set a "sleep" expression (millisecond delay) that uses a synchronized last-execute-time check (lines 119-141) to pace requests. There is no built-in browser rendering, no CSS pre-loading, and no AJAX-wait logic.

How is content extracted or converted?

answered

Extraction is split between two layers. ExtractFunctionExecutor (spider-flow-core/.../extract/ExtractFunctionExecutor.java) exposes static extraction functions callable in flow expressions: extract.jsonpath(root, '$.path') uses FastJson's JSONPath; extract.xpath(element, '//xpath') and extract.xpaths() use the Xsoup library on Jsoup-parsed Elements; extract.regx(content, pattern) and extract.regxs() use Java Matcher with DOTALL mode; extract.selector() and extract.selectors() use Jsoup's CSS selector engine, supporting return types: "text", "html", "outerhtml", "attr", "element". All delegate to ExtractUtils (spider-flow-core/.../utils/ExtractUtils.java) which caches compiled Regex patterns and provides helpers for each strategy. The ResponseFunctionExtension provides convenience shorthands directly on the resp variable: resp.selector(), resp.xpath(), resp.regx(), resp.jsonpath(), resp.links(), and resp.images(). There is no readability/boilerplate removal (no Readability or similar heuristic). Extraction is purely selector-driven. For output, the OutputExecutor shape (spider-flow-core/.../OutputExecutor.java) can write extracted fields to a database (via Spring JdbcTemplate, configurable data source) or to CSV files (Apache Commons CSV). It also supports outputting all flow variables. The ExecuteSQLExecutor shape can run arbitrary SQL (select/insert/update/delete), supporting parameter binding and batch operations.

How are LLMs used, if at all?

not applicable

This repository does not use LLMs in any form. There are zero imports or references to any LLM provider (OpenAI, Anthropic, Google Generative AI, Mistral, Cohere, Ollama) across the entire Java codebase. A grep for 'openai', 'gpt', 'llm', 'chatgpt', 'embedding', 'vector', 'nlp', or 'natural language' returned no relevant matches in the core modules - only hits in third-party JS libraries and Java AST utility method names that happen to match substrings. The project operates entirely on rule-based extraction: CSS selectors, XPath, regex, and JSONPath expressions defined by the user in the visual flow editor. There is no automatic summarization, no AI-based field extraction, no chunking of large pages for LLM consumption, no structured output via JSON schema prompts, and no cost controls related to AI model usage. The extraction paradigm is strictly declarative selectors + programmatic transformation functions (string, date, list, math operations) available through the expression engine.

How are anti-bot measures, proxies and fingerprinting handled?

answered

Anti-bot measures are rudimentary and user-driven. Proxy support exists per request node: the "proxy" field accepts an expression that resolves to a host:port string, which is parsed and set on the Jsoup Connection via request.proxy(host, port) (RequestExecutor lines 241-256). User-agent spoofing is minimal: the only hardcoded UA is in the FileUtils.downloadFile() helper (spider-flow-core/.../utils/FileUtils.java line 211), which sets Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:68.0) Gecko/20100101 Firefox/68.0 for file downloads only. Custom headers (including User-Agent) can be set per request node through the header name/value pairs. Cookie management is comprehensive: the CookieContext (spider-flow-api/.../context/CookieContext.java, extends HashMap) auto-stores cookies from responses and auto-attaches them to subsequent requests when cookie-auto-set is enabled (RequestExecutor lines 203-208). Rate limiting is manual: each node can define a "sleep" expression that enforces a minimum time between requests to the same node, using a last-execute-time check (lines 119-141). Retry logic is built in: configurable retry count and interval per node (lines 145-148, 290-293). CAPTCHA handling is absent. Fingerprint spoofing beyond custom headers is absent. TLS validation can be disabled per node (lines 187-190). Automated proxy rotation, stealth browser patches, and request fingerprint randomization are not implemented. The README mentions a separate proxypool plugin (spider-flow-proxypool) and an OCR plugin (spider-flow-ocr) available as external add-ons.

How is crawling at scale implemented?

answered

Crawling is structured as a directed-acyclic-graph (DAG) workflow rather than traditional BFS/DFS queues. The Spider class (spider-flow-core/.../Spider.java) is the core engine. It deserializes the flow XML into a SpiderNode tree. Execution starts at the root node and fans out to connected children. Concurrency is controlled by SpiderFlowThreadPoolExecutor with a configurable global cap (default 64 threads via spider.thread.max) and per-flow cap (default 8 via spider.thread.default, Spider.java lines 47-51). Each flow gets a SubThreadPoolExecutor with a submit strategy - one of: Random, Linked (depth-first), ChildPrior, or ParentPrior (lines 112-123). Nodes that return isThread()=true execute asynchronously in the pool; synchronous nodes (e.g., ForkJoin) run inline. URL dedup uses a Bloom filter (Guava library) persisted to disk per flow ID (RequestExecutor lines 444-480), configurable via spider.bloomfilter.capacity and spider.bloomfilter.error-rate. Loop/pagination is handled by the LoopExecutor combined with the node-level loop count/collection expression in executeNode() (Spider.java lines 218-253). Scheduling uses Quartz via SpiderJob (spider-flow-core/.../job/SpiderJob.java) which runs as a QuartzJobBean with cron expression support, persisting Task records. SpiderJobManager manages job lifecycle (add/remove). On app start, all enabled cron flows are registered (SpiderFlowService.initJobs(), lines 54-71). There are no distributed workers, no external message queues, no crawl frontier/middleware architecture, and no robots.txt parser. The project uses its own custom thread pool with a submit-scheduler pattern rather than standard executor services. ProcessExecutor allows nesting one flow as a sub-routine of another.

Editor's note. Correction: URL dedup is opt-in per request node (the repeat-enable flag); without it no Bloom filter is created. Scheduled cron runs are disabled by the shipped application.properties (spider.job.enable=false), and flows are not limited to DAGs; cycles are allowed and only test runs have dead-cycle detection.

What is the developer interface?

answered

The primary interface is a browser-based visual flow editor using mxGraph (a JavaScript diagramming library) for drag-and-drop DAG construction. The editor (editor.html) provides a toolbar for save, test, debug (step-through with WebSocket), undo/redo, and XML editing. Each shape node gets a context panel where users configure request URLs, headers, extraction rules (CodeMirror-based expression editors with auto-complete), and output mappings. The REST API is exposed through several Spring controllers: SpiderRestController (/rest/run/{id} for synchronous execution returning JSON outputs; /rest/runAsync/{id} for async with a task ID; /rest/stop/{taskId}; /rest/status/{taskId}). SpiderFlowController (/spider/) manages CRUD for flow definitions (list, save, get, remove, copy, start/stop cron, log download/view, list shapes/grammars/plugin configs). TaskController (/task/) manages execution history and task stopping. DataSourceController (/datasource/) manages JDBC data source connections. A WebSocket endpoint (/ws) enables real-time flow testing and step-by-step debugging (WebSocketEditorServer.java). The output formats are: (1) JSON returned from the REST sync endpoint as SpiderOutput[] (name/value pairs per node), (2) database tables via the OutputExecutor's JDBC output, (3) CSV files written to disk. Configuration is via application.properties (workspace path, thread limits, Bloom filter settings, mail, DB credentials). The frontend (separate repo spider-flow-vue) uses LayUI. The base URL is from the Gitee-hosted repository. Language bindings: Java only; no Python, Node.js, or SDK bindings exist.

Editor's note. Correction: the UI served by this repo is the static LayUI + mxGraph editor in spider-flow-web/src/main/resources/static; spider-flow-vue is an optional separate frontend. The web app has no authentication at all, so the REST, WebSocket and SQL-executing endpoints are open to anyone who can reach port 8088.