# ssssssss-team/spider-flow

> Java/Spring Boot visual flowchart crawler builder on Jsoup and selector expressions; no JS rendering or LLMs, last commit 2022.

- Category: [AI web scraping](https://llms-technical-reviews.com/ai-scraping/)
- Repository: https://github.com/ssssssss-team/spider-flow (reviewed at commit `c799cca99c7d064673dc79e6ef06a08bbc0e292d`, 2022-10-25)
- Stars: 11358 · Language: Java · License: MIT
- Canonical page: https://llms-technical-reviews.com/p/spider-flow/

## Overview

Spider-Flow is a **visual crawler builder**, not an AI scraper. It is a Spring Boot 2.0 web app (Java 8). You draw a scraping job as a flowchart in the browser: a start node, request nodes, variable and loop nodes, SQL and output nodes. Arrows between nodes carry conditions. Each node holds small expressions such as `${extract.xpath(resp.element(), '//h1')}`. The server runs the chart on a thread pool, fetches pages with Jsoup, and writes rows to a database or CSV. It has no LLM code and no headless browser.

The project's last commit on the default branch (the one reviewed here) is from October 2022, version 0.5.0. Dependencies are from that period: Spring Boot 2.0.7, Nashorn for user-defined JavaScript functions (removed from the JDK in Java 15), and FastJson. Selenium rendering, proxy pools, OCR, Redis and Mongo outputs are separate plugins in other Gitee repositories. They are not in this codebase.

**Where it fits in an AI scraping stack:** it is a low-code fetch-and-select tool for simple, server-rendered sites, used by people who prefer a flowchart to code. It is a weak fit for modern AI pipelines. It cannot render JavaScript without the external plugin. It has no HTML-to-Markdown step, so you cannot easily send clean text to a model. The UI has no authentication. If you only need a scheduler around scripts, Crawlab is the closer fit. If you need LLM-ready output, use a newer crawler.

## Architecture

```mermaid
flowchart LR
  ED["Browser editor (mxGraph)"] -->|"flow XML"| WEB["Spring MVC controllers"]
  ED -->|"WebSocket /ws test/debug"| WS["WebSocketEditorServer"]
  WEB --> DB["MySQL (flows, tasks, data sources)"]
  CRON["Quartz SpiderJob"] --> SP["Spider engine"]
  WEB -->|"/rest/run"| SP
  WS --> SP
  SP --> POOL["Thread pool + submit strategy"]
  POOL --> EX["Shape executors"]
  EX --> REQ["RequestExecutor (Jsoup)"]
  EX --> OUT["Output / SQL executors"]
  EX --> EXPR["Expression engine + functions"]
  OUT --> RES["JDBC tables / CSV"]
```

| Component | Path | Role |
|---|---|---|
| API module | `spider-flow-api/` | Interfaces: `ShapeExecutor`, `FunctionExecutor`, `FunctionExtension`, `PluginConfig`, `SpiderContext` |
| Engine | `spider-flow-core/.../core/Spider.java` | Parses flow XML, walks nodes, runs loops, manages futures |
| Thread pool | `spider-flow-api/.../concurrent/` | Global pool with per-flow sub-pools and four submit strategies |
| Shape executors | `spider-flow-core/.../executor/shape/` | `RequestExecutor`, `VariableExecutor`, `LoopExecutor`, `ForkJoinExecutor`, `OutputExecutor`, `ExecuteSQLExecutor`, `ProcessExecutor`, and more |
| Functions | `spider-flow-core/.../executor/function/` | `extract`, `string`, `date`, `json`, `file`, `md5` and other expression prefixes, plus type extensions (`resp.xpath`, `element.selector`) |
| Expressions | `spider-flow-core/.../expression/` | Template engine with its own tokenizer, parser and reflection-based interpreter |
| HTTP | `spider-flow-core/.../io/HttpRequest.java` | Thin wrapper over a Jsoup `Connection` |
| Scheduling | `spider-flow-core/.../job/` | Quartz `SpiderJob` and `SpiderJobManager` |
| Web | `spider-flow-web/` | Controllers, WebSocket debugger, static LayUI + mxGraph UI, `application.properties` |

## How a request flows

Take a flow "start → request list page → loop over links → request detail → output":

1. **Trigger.** A cron fire (`SpiderJob.executeInternal`), a REST call (`/rest/run/{id}` or `/rest/runAsync/{id}`) or the editor's Test button (WebSocket `test`/`debug` event) calls `Spider.run` or `runWithTest` ([SpiderJob.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/job/SpiderJob.java#L50-L94), [SpiderRestController.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-web/src/main/java/org/spiderflow/controller/SpiderRestController.java#L102-L120), [WebSocketEditorServer.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-web/src/main/java/org/spiderflow/websocket/WebSocketEditorServer.java#L31-L54)).
2. **Load and pool.** `run` parses the stored mxGraph XML into a `SpiderNode` tree. `executeRoot` reads the start node's thread count (default 8) and submit strategy (random, linked, child-first or parent-first). It then creates a sub-pool of the global 64-thread executor ([Spider.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/Spider.java#L70-L127)).
3. **Schedule loop.** One coordinator task runs the root, then polls the future queue every millisecond. When a future is done (picked by the strategy's comparator), it calls `allowExecuteNext` and then `executeNextNodes` on the children ([Spider.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/Spider.java#L132-L184)).
4. **Execute a node.** `executeNode` checks the arrow's condition and exception-flow flag. It evaluates the node's loop count or collection and creates one task per iteration. Each task gets a copy of the variables, with the index and `item` added. Async shapes go to the pool, and sync shapes run inline. A thrown exception is stored as `ex`, so later arrows can branch on failure ([Spider.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/Spider.java#L201-L308)).
5. **Request.** `RequestExecutor.execute` applies the per-node sleep and checks the optional Bloom filter. It builds the URL, method, headers, cookies, body and `host:port` proxy from expressions, then calls `HttpRequest.execute`. That call uses Jsoup with `ignoreContentType`, `ignoreHttpErrors` and an unlimited body size. Only HTTP 200 counts as success. Anything else is retried up to the configured count. On success the response is stored as the `resp` variable ([RequestExecutor.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L116-L300), [HttpRequest.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/io/HttpRequest.java#L121-L128)).
6. **Extract.** Variable and output nodes evaluate expressions such as `${extract.selector(resp.html,'div > a','attr','href')}` or `${resp.jsonpath('$.data')}`. These call Jsoup CSS selectors, Xsoup XPath, regex or FastJson JSONPath ([ExtractFunctionExecutor.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/function/ExtractFunctionExecutor.java#L88-L130)).
7. **Output.** `OutputExecutor` evaluates each name/value pair. It adds the pairs to the run's outputs (returned as JSON by `/rest/run`) and can insert a row into a configured JDBC data source or append to a CSV ([OutputExecutor.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/OutputExecutor.java#L60-L105)).

## Key components

### The flow engine

Flows are graphs, not strict DAGs. A node can loop back, so test runs count executions and stop past `spider.detect.dead-cycle` (default 5000) ([Spider.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/Spider.java#L87-L104)). Production runs have no such guard. The coordinator's 1 ms polling loop scans the whole future queue on each pass. That is fine for hundreds of pending requests and wasteful for many thousands. A `ForkJoin` shape uses per-node counters to wait for all branches. `ProcessExecutor` runs another flow as a sub-routine.

### Request node

All fetching is plain HTTP through Jsoup. Each node has timeout (default 60 s), method, redirect and TLS-validation toggles, form or raw bodies, file uploads, a response charset override, and retry count and interval. Cookies set by responses go into a shared `CookieContext` and are re-sent by default. Politeness is manual: a per-node `sleep` expression enforces a minimum gap between that node's requests. There is no robots.txt handling, user-agent rotation or proxy pool. Proxy strings without credentials are the only proxy option.

### URL dedup

Dedup is opt-in per request node. When it is on, `createBloomFilter` loads `<workspace>/<flowId>/url.bf` or creates a Guava Bloom filter (capacity 1,000,000, error rate 0.0001 by default). `afterEnd` writes it back to disk, so dedup carries across runs ([RequestExecutor.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L444-L480)). A URL is recorded only after a 200 response, so failed URLs are retried on the next run.

### Expression language

`DefaultExpressionEngine` builds a context with every `FunctionExecutor` under its prefix and every global variable. It then renders the `${...}` template ([DefaultExpressionEngine.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/expression/DefaultExpressionEngine.java#L32-L50)). The interpreter calls methods by reflection. `FunctionExtension` classes add methods to existing types, such as `resp.xpath(...)` or `list.join(...)`. User-defined functions are JavaScript, compiled by Nashorn through `ScriptManager`.

## Extending it

- **New node types.** Implement `ShapeExecutor` (`supportShape`, `execute`, optional `allowExecuteNext`/`isThread`) as a Spring bean. `ExecutorsUtils` finds it by shape name ([ShapeExecutor.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-api/src/main/java/org/spiderflow/executor/ShapeExecutor.java#L13-L46)).
- **New functions.** Implement `FunctionExecutor` with a prefix and public static methods. `@Comment` and `@Example` feed the editor's autocomplete.
- **Plugins.** The external Selenium, proxy-pool, OCR, Redis, OSS, Mongo and mailbox plugins are jars that add executors and register a `PluginConfig`.
- **Listeners.** `SpiderListener.beforeStart`/`afterEnd` hooks run around each flow. The Bloom filter persistence uses them.

## Running it

- Build with Maven (`spider-flow-web` produces `spider-flow.jar`). Import `db/spiderflow.sql` into MySQL and set the datasource in `application.properties`. The shipped values are `root` / `123456789` on `localhost:3306`. The app listens on 8088. The Dockerfile uses `java:8`.
- Scheduled flows do nothing until you set `spider.job.enable=true`. The shipped properties file sets it to `false`, and `SpiderJob.executeInternal` returns early ([SpiderJob.java](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/job/SpiderJob.java#L40-L58)).
- Use a Java 8–14 runtime. Custom JavaScript functions need Nashorn.
- It is a single process. There are no distributed workers, and all concurrency is threads inside one JVM.

## Strengths and caveats

- **Strength: approachable.** A non-programmer can build a paginated list → detail → database job in the editor, test it step by step over WebSocket and schedule it.
- **Strength: useful primitives.** XPath, CSS, regex and JSONPath in one expression language, plus loop/fork-join nodes, cookie carry-over and persistent Bloom-filter dedup.
- **Caveat: not AI, no rendering.** It has no LLM calls, no Markdown conversion and no JS execution without the external Selenium plugin. Its output is selector-defined fields only.
- **Caveat: no authentication.** There is no login and no security filter. Anyone who reaches the port can run flows, use the WebSocket test endpoint, or run SQL against configured data sources through `ExecuteSQLExecutor`. Keep it on a private network.
- **Caveat: unmaintained stack.** Spring Boot 2.0, Java 8, Nashorn and an old FastJson make upgrades and security patching your job.
- **Caveat: brittle success rule.** Only status 200 counts as success. A 201, 204 or unfollowed 3xx is retried and then logged as a failure.

*Sources: code at c799cca, deepwiki-open wiki (13 pages), OpenDeepWiki wiki (12 pages), verified Q&A.*

## How ssssssss-team/spider-flow answers the AI web scraping questions

### How are pages fetched and rendered? (answered)

All HTTP requests go through the **HttpRequest** class (spider-flow-core/.../io/HttpRequest.java), which wraps **Jsoup's Connection** - a pure-Java HTTP client, not a headless browser. The **RequestExecutor** shape (spider-flow-core/.../RequestExecutor.java) configures each request: method (GET/POST etc.), URL (via expression evaluation), timeout (default 60s, "timeout" field), follow-redirect toggle, TLS-certificate validation toggle, custom headers, and per-node/project-level cookies. JS rendering is **not built in**; the README links a separate Selenium plugin (spider-flow-selenium). The Jsoup connection calls `ignoreContentType(true)` and `ignoreHttpErrors(true)` with `maxBodySize(0)` (unlimited) in HttpRequest.execute() (lines 121-128). The response is wrapped as **HttpResponse** (spider-flow-core/.../io/HttpResponse.java) implementing the **SpiderResponse** interface. It exposes `getHtml()` (raw body), `getBytes()` (byte array), `getStream()` (InputStream), `getJson()` (parsed via FastJson), `getContentType()`, and `getStatusCode()` - so all content types (HTML, JSON, binary) are captured. A `setCharset()` override is available per node (RequestExecutor line 267-271). The waiting/politeness strategy is manual: each request node can set a "sleep" expression (millisecond delay) that uses a synchronized last-execute-time check (lines 119-141) to pace requests. There is no built-in browser rendering, no CSS pre-loading, and no AJAX-wait logic.


Citations: [spider-flow-core/src/main/java/org/spiderflow/core/io/HttpRequest.java:17-129](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/io/HttpRequest.java#L17-L129) · [spider-flow-core/src/main/java/org/spiderflow/core/io/HttpResponse.java:16-104](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/io/HttpResponse.java#L16-L104) · [spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java:117-317](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L117-L317) · [spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java:88-93](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L88-L93)

### How is content extracted or converted? (answered)

Extraction is split between two layers. **ExtractFunctionExecutor** (spider-flow-core/.../extract/ExtractFunctionExecutor.java) exposes static extraction functions callable in flow expressions: `extract.jsonpath(root, '$.path')` uses FastJson's JSONPath; `extract.xpath(element, '//xpath')` and `extract.xpaths()` use the **Xsoup** library on Jsoup-parsed Elements; `extract.regx(content, pattern)` and `extract.regxs()` use Java `Matcher` with DOTALL mode; `extract.selector()` and `extract.selectors()` use **Jsoup's CSS selector** engine, supporting return types: "text", "html", "outerhtml", "attr", "element". All delegate to **ExtractUtils** (spider-flow-core/.../utils/ExtractUtils.java) which caches compiled Regex patterns and provides helpers for each strategy. The **ResponseFunctionExtension** provides convenience shorthands directly on the `resp` variable: `resp.selector()`, `resp.xpath()`, `resp.regx()`, `resp.jsonpath()`, `resp.links()`, and `resp.images()`. There is no readability/boilerplate removal (no Readability or similar heuristic). Extraction is purely selector-driven. For output, the **OutputExecutor** shape (spider-flow-core/.../OutputExecutor.java) can write extracted fields to a **database** (via Spring JdbcTemplate, configurable data source) or to **CSV files** (Apache Commons CSV). It also supports outputting all flow variables. The **ExecuteSQLExecutor** shape can run arbitrary SQL (select/insert/update/delete), supporting parameter binding and batch operations.


Citations: [spider-flow-core/src/main/java/org/spiderflow/core/executor/function/ExtractFunctionExecutor.java:16-150](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/function/ExtractFunctionExecutor.java#L16-L150) · [spider-flow-core/src/main/java/org/spiderflow/core/utils/ExtractUtils.java:21-180](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/utils/ExtractUtils.java#L21-L180) · [spider-flow-core/src/main/java/org/spiderflow/core/executor/function/extension/ResponseFunctionExtension.java:19-127](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/function/extension/ResponseFunctionExtension.java#L19-L127)

### How are LLMs used, if at all? (not applicable)

This repository does not use LLMs in any form. There are zero imports or references to any LLM provider (OpenAI, Anthropic, Google Generative AI, Mistral, Cohere, Ollama) across the entire Java codebase. A grep for 'openai', 'gpt', 'llm', 'chatgpt', 'embedding', 'vector', 'nlp', or 'natural language' returned no relevant matches in the core modules - only hits in third-party JS libraries and Java AST utility method names that happen to match substrings. The project operates entirely on rule-based extraction: CSS selectors, XPath, regex, and JSONPath expressions defined by the user in the visual flow editor. There is no automatic summarization, no AI-based field extraction, no chunking of large pages for LLM consumption, no structured output via JSON schema prompts, and no cost controls related to AI model usage. The extraction paradigm is strictly declarative selectors + programmatic transformation functions (string, date, list, math operations) available through the expression engine.


Citations: [spider-flow-core/src/main/java/org/spiderflow/core/executor/function/ExtractFunctionExecutor.java:14-150](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/function/ExtractFunctionExecutor.java#L14-L150) · [spider-flow-core/src/main/java/org/spiderflow/core/utils/ExtractUtils.java:21-180](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/utils/ExtractUtils.java#L21-L180)

### How are anti-bot measures, proxies and fingerprinting handled? (answered)

Anti-bot measures are rudimentary and user-driven. **Proxy support** exists per request node: the "proxy" field accepts an expression that resolves to a `host:port` string, which is parsed and set on the Jsoup Connection via `request.proxy(host, port)` (RequestExecutor lines 241-256). **User-agent spoofing** is minimal: the only hardcoded UA is in the `FileUtils.downloadFile()` helper (spider-flow-core/.../utils/FileUtils.java line 211), which sets `Mozilla/5.0 (Windows NT 10.0; Win64; x64; rv:68.0) Gecko/20100101 Firefox/68.0` for file downloads only. Custom headers (including User-Agent) can be set per request node through the header name/value pairs. **Cookie management** is comprehensive: the CookieContext (spider-flow-api/.../context/CookieContext.java, extends HashMap) auto-stores cookies from responses and auto-attaches them to subsequent requests when `cookie-auto-set` is enabled (RequestExecutor lines 203-208). **Rate limiting** is manual: each node can define a "sleep" expression that enforces a minimum time between requests to the same node, using a last-execute-time check (lines 119-141). **Retry logic** is built in: configurable retry count and interval per node (lines 145-148, 290-293). **CAPTCHA handling** is absent. **Fingerprint spoofing** beyond custom headers is absent. **TLS validation** can be disabled per node (lines 187-190). Automated proxy rotation, stealth browser patches, and request fingerprint randomization are not implemented. The README mentions a separate proxypool plugin (spider-flow-proxypool) and an OCR plugin (spider-flow-ocr) available as external add-ons.


Citations: [spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java:241-256](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L241-L256) · [spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java:119-141](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L119-L141) · [spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java:145-148](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L145-L148) · [spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java:186-190](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L186-L190) · [spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java:203-208](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L203-L208) · [spider-flow-core/src/main/java/org/spiderflow/core/utils/FileUtils.java:179-212](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/utils/FileUtils.java#L179-L212)

### How is crawling at scale implemented? (answered)

Crawling is structured as a **directed-acyclic-graph (DAG) workflow** rather than traditional BFS/DFS queues. The **Spider** class (spider-flow-core/.../Spider.java) is the core engine. It deserializes the flow XML into a **SpiderNode** tree. Execution starts at the root node and fans out to connected children. Concurrency is controlled by `SpiderFlowThreadPoolExecutor` with a configurable global cap (default 64 threads via `spider.thread.max`) and per-flow cap (default 8 via `spider.thread.default`, Spider.java lines 47-51). Each flow gets a `SubThreadPoolExecutor` with a **submit strategy** - one of: Random, Linked (depth-first), ChildPrior, or ParentPrior (lines 112-123). Nodes that return `isThread()=true` execute asynchronously in the pool; synchronous nodes (e.g., ForkJoin) run inline. **URL dedup** uses a **Bloom filter** (Guava library) persisted to disk per flow ID (RequestExecutor lines 444-480), configurable via `spider.bloomfilter.capacity` and `spider.bloomfilter.error-rate`. **Loop/pagination** is handled by the LoopExecutor combined with the node-level loop count/collection expression in `executeNode()` (Spider.java lines 218-253). **Scheduling** uses **Quartz** via **SpiderJob** (spider-flow-core/.../job/SpiderJob.java) which runs as a QuartzJobBean with cron expression support, persisting Task records. **SpiderJobManager** manages job lifecycle (add/remove). On app start, all enabled cron flows are registered (SpiderFlowService.initJobs(), lines 54-71). There are **no distributed workers**, **no external message queues**, no crawl frontier/middleware architecture, and no robots.txt parser. The project uses its own custom thread pool with a submit-scheduler pattern rather than standard executor services. ProcessExecutor allows nesting one flow as a sub-routine of another.

> **Editor's note.** Correction: URL dedup is opt-in per request node (the repeat-enable flag); without it no Bloom filter is created. Scheduled cron runs are disabled by the shipped application.properties (`spider.job.enable=false`), and flows are not limited to DAGs; cycles are allowed and only test runs have dead-cycle detection.

Citations: [spider-flow-core/src/main/java/org/spiderflow/core/Spider.java:47-68](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/Spider.java#L47-L68) · [spider-flow-core/src/main/java/org/spiderflow/core/Spider.java:109-308](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/Spider.java#L109-L308) · [spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java:444-480](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/executor/shape/RequestExecutor.java#L444-L480) · [spider-flow-core/src/main/java/org/spiderflow/core/job/SpiderJob.java:30-102](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/job/SpiderJob.java#L30-L102) · [spider-flow-core/src/main/java/org/spiderflow/core/job/SpiderJobManager.java:26-89](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-core/src/main/java/org/spiderflow/core/job/SpiderJobManager.java#L26-L89) · [spider-flow-api/src/main/java/org/spiderflow/concurrent/SpiderFlowThreadPoolExecutor.java:8-201](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-api/src/main/java/org/spiderflow/concurrent/SpiderFlowThreadPoolExecutor.java#L8-L201)

### What is the developer interface? (answered)

The primary interface is a **browser-based visual flow editor** using mxGraph (a JavaScript diagramming library) for drag-and-drop DAG construction. The editor (editor.html) provides a toolbar for save, test, debug (step-through with WebSocket), undo/redo, and XML editing. Each shape node gets a context panel where users configure request URLs, headers, extraction rules (CodeMirror-based expression editors with auto-complete), and output mappings. The **REST API** is exposed through several Spring controllers: **SpiderRestController** (`/rest/run/{id}` for synchronous execution returning JSON outputs; `/rest/runAsync/{id}` for async with a task ID; `/rest/stop/{taskId}`; `/rest/status/{taskId}`). **SpiderFlowController** (`/spider/`) manages CRUD for flow definitions (list, save, get, remove, copy, start/stop cron, log download/view, list shapes/grammars/plugin configs). **TaskController** (`/task/`) manages execution history and task stopping. **DataSourceController** (`/datasource/`) manages JDBC data source connections. A **WebSocket endpoint** (`/ws`) enables real-time flow testing and step-by-step debugging (WebSocketEditorServer.java). The **output formats** are: (1) JSON returned from the REST sync endpoint as `SpiderOutput[]` (name/value pairs per node), (2) database tables via the OutputExecutor's JDBC output, (3) CSV files written to disk. **Configuration** is via `application.properties` (workspace path, thread limits, Bloom filter settings, mail, DB credentials). The frontend (separate repo spider-flow-vue) uses LayUI. The base URL is from the Gitee-hosted repository. Language bindings: Java only; no Python, Node.js, or SDK bindings exist.

> **Editor's note.** Correction: the UI served by this repo is the static LayUI + mxGraph editor in spider-flow-web/src/main/resources/static; spider-flow-vue is an optional separate frontend. The web app has no authentication at all, so the REST, WebSocket and SQL-executing endpoints are open to anyone who can reach port 8088.

Citations: [spider-flow-web/src/main/resources/static/editor.html:1-63](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-web/src/main/resources/static/editor.html#L1-L63) · [spider-flow-web/src/main/java/org/spiderflow/websocket/WebSocketEditorServer.java:21-60](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-web/src/main/java/org/spiderflow/websocket/WebSocketEditorServer.java#L21-L60) · [spider-flow-web/src/main/java/org/spiderflow/controller/SpiderRestController.java:28-122](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-web/src/main/java/org/spiderflow/controller/SpiderRestController.java#L28-L122) · [spider-flow-web/src/main/java/org/spiderflow/controller/SpiderFlowController.java:48-220](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-web/src/main/java/org/spiderflow/controller/SpiderFlowController.java#L48-L220) · [spider-flow-web/src/main/resources/static/index.html:1-60](https://github.com/ssssssss-team/spider-flow/blob/c799cca99c7d064673dc79e6ef06a08bbc0e292d/spider-flow-web/src/main/resources/static/index.html#L1-L60)
