crawlab-team/crawlab
Go and MongoDB platform that schedules and runs your own spider programs across nodes; it does no fetching or extraction itself.
Overview
Crawlab is not a scraper and has no AI features. It is a crawler management platform: a Go server, a MongoDB database and a web UI that run your spider programs on one or more machines. A spider can be a Scrapy project, a Node/Puppeteer script, a Go binary or any other command. Crawlab stores the code, starts it as an OS subprocess on a node, captures its logs and collects the records it reports back. Scheduling is done with cron expressions. Nothing in the repository fetches a web page, parses HTML or calls a model.
Where it fits in an AI scraping stack: it is the job-control layer. Your spider (perhaps a Crawl4AI or Firecrawl client, or a Playwright script that sends Markdown to an LLM) is the payload. Crawlab gives that payload a UI, cron schedules, per-node concurrency limits, live logs and a results table. Anything about fetching, stealth or extraction must live in the spider.
The pinned commit (October 2024) is the v0.6 rewrite. The code has *V2 services next to legacy ones, and some features are only half moved over. The sections below note where that matters.
Architecture
flowchart LR
UI["Web UI (crawlab-ui package)"] -->|"/api via nginx"| API["Gin REST API"]
API --> ADM["Spider admin service"]
SCH["Cron schedule service"] --> ADM
ADM --> Q["Mongo task queue"]
API --> DB["MongoDB"]
W["Worker task handler"] -->|"gRPC Fetch"| GS["Master gRPC server"]
GS --> Q
W --> R["Task runner"]
R -->|"HTTP /sync"| API
R --> P["Spider subprocess"]
P -->|"gRPC stream: data, logs"| GS
GS --> DB
M["Master monitor"] -->|"ping every 15 s"| W
| Component | Path | Role |
|---|---|---|
| Entry point | core/cmd/server.go, core/apps/ |
crawlab server starts the API, gRPC server and node services |
| REST API | core/controllers/router_v2.go |
Gin routes for spiders, tasks, schedules, nodes, results, export, sync |
| Spider admin | core/spider/admin/service_v2.go |
Turns “run spider” into a task and a queue item |
| Task scheduler | core/task/scheduler/service_v2.go |
Enqueue, cancel, status recovery, 30-day cleanup |
| Task handler | core/task/handler/service_v2.go |
Per-node loop that fetches tasks and runs them |
| Task runner | core/task/handler/runner_v2.go |
Syncs files, builds the command, sets env, starts and watches the process |
| gRPC task server | core/grpc/server/task_server_v2.go |
Dequeues tasks; receives data and log streams |
| Master service | core/node/service/master_service_v2.go |
Monitors workers, marks them offline |
| Result writer | core/task/stats/service_v2.go |
Writes reported records to the spider’s collection |
| Schedules | core/schedule/service_v2.go |
robfig/cron entries that call the spider admin |
| Frontend | frontend/, nginx/crawlab.conf |
Loads the prebuilt crawlab-ui npm package; nginx serves it on 8080 |
How a request flows
Take a click on “Run” for a spider:
- API.
PostSpiderRunbindsSpiderRunOptions(mode, node ids, cmd, param, priority) and calls the spider admin service (spider_v2.go). - Task creation.
scheduleTasksbuilds oneTaskV2and fills empty fields from the spider’s defaults. If the mode named any nodes, it pins the task to the first one (service_v2.go).Enqueueinserts the task, aTaskQueueItemV2with the priority and node id, and an empty stats row (scheduler/service_v2.go). - Fetch. On every node,
ServiceV2.Fetchticks. It skips inactive or disabled nodes and nodes already atMaxRunners(default 8). Otherwise it calls the master’s gRPCFetch(handler/service_v2.go). The master first looks for a queue item assigned to that node, then for an unassigned one. It sorts by priority and id, sets the task’s node and deletes the queue item (task_server_v2.go, L242-L263). - Sync code. On a worker,
RunnerV2.syncFilesasks the master’s/sync/:id/scanfor a file list with hashes, deletes local files that are gone, and downloads changed files with 10 parallel requests (runner_v2.go). - Build and start.
configureCmdjoins the command and param and passes them tosys_exec.BuildCmd. That function splits the string on single spaces, with no shell and no quoting (runner_v2.go, sys_exec_linux.go).configureEnvaddsCRAWLAB_TASK_ID, the gRPC address and auth key, and every global environment variable from the database (runner_v2.go).Runstarts the process, pipes stdout and stderr to the log stream, and waits on a signal channel for finish, cancel, error or “lost” (runner_v2.go). - Report data. The spider uses the separate Crawlab SDK (
save_item, or a Scrapy pipeline) to stream records over gRPC with the task id from the environment.SubscribesendsINSERT_DATAtohandleInsertDataandINSERT_LOGSto the log driver (task_server_v2.go, L212-L232).InsertDatawrites the batch into the spider’s Mongo collection and updates the result count (stats/service_v2.go).
Key components
Task queue
The queue is the task_queue_items Mongo collection, polled by every node through the master. This design is simple and needs no broker. Two details matter at scale. The dequeue is a find followed by a delete. It runs inside RunTransactionWithContext, but the model-service calls do not use the session context, so the transaction does not protect them. Two nodes polling at the same moment can claim the same item. And polling is the only dispatch: a task waits up to one fetch interval before a node picks it up.
Run modes
The UI offers “all nodes”, “random” and “selected nodes”. In the V2 admin service, all-nodes and selected-nodes both produce one task, assigned to the first node in the list. The legacy service created one task per node; that fan-out was not ported (legacy service.go). If you need the same job on every worker, schedule one task per node.
Node monitoring
The master runs monitor every 15 s. For each worker it subscribes, pings over gRPC and updates the free runner count. A worker that fails either check is marked offline (master_service_v2.go). The runner also has a process health check, but it returns at once when cmd.ProcessState is nil, and that is always the case while the process runs. So it never reports a lost process (runner_v2.go).
Results
Records go to the spider’s col_name collection in MongoDB. Only Pro editions (edition: global.edition.pro) route them to another database through a registry. The core/ds folder has MySQL, PostgreSQL, Elasticsearch, Kafka and other drivers, but no package imports it. Watch for one bug in the stats service. On a cache miss, getDatabaseServiceItem returns the nil item it looked up, not the entry it just cached. The first InsertData call for each task then dereferences nil (stats/service_v2.go). Results can be exported as CSV or JSON from the API.
Schedules
schedule/service_v2.go holds a robfig/cron instance. Each entry loads the schedule and the spider, merges their mode, nodes, cmd and param, and calls the same Schedule path as the Run button (schedule/service_v2.go).
Extending it
- Any language. A spider is a folder plus a
cmdstring. Install the runtime on the node image. Global environment variables reach every task, which is a good place for API keys and proxy URLs. - Crawlab SDK. Spiders report records with
save_item(Python) orCrawlabPipeline(Scrapy). Without the SDK, the platform still captures logs but sees no data. - Git-backed spiders.
SpiderV2.GitIdandGitRootPathsync code from a repository (vcs/module). - REST API. Every UI action is a Gin route under
/api, so CI or an agent can create spiders, start runs and read results.
Running it
- Use Docker Compose: one
crawlabteam/crawlabmaster (CRAWLAB_NODE_MASTER=Y), any number of workers pointing at it throughCRAWLAB_GRPC_ADDRESS, and MongoDB. The UI is on port 8080. nginx forwards/api/to the Go server on 8000. gRPC uses 9666. - The images are built from separate backend, frontend and plugin images. This repository’s
frontend/folder only bootstraps thecrawlab-ui0.6.3 npm package (main.ts). - Change the defaults. The gRPC auth key falls back to the constant
Crawlab2021!. The/sync/:id/scanand/sync/:id/downloadroutes are in the anonymous route group, and they join thepathquery onto the workspace directory without checking that it stays inside it (router_v2.go, sync_v2.go). Keep the API on a private network.
Strengths and caveats
- Strength: framework-neutral. It runs anything with a command line, so it can schedule AI scrapers written in any stack without code changes.
- Strength: operations UI. Live logs, task history, cron schedules, per-node runner caps and a data browser come ready-made.
- Strength: few moving parts. Only MongoDB and the Crawlab binary. Workers sync code from the master automatically.
- Caveat: no scraping features. It has no fetching, proxies, stealth, robots.txt, URL dedup, extraction or LLM support. Crawl-level concerns stay inside each spider.
- Caveat: V2 migration gaps. Run modes do not fan out, the health check is dead code, the
dsdrivers are unused, and the stats cache has a nil bug. The two-version code makes it hard to tell what is live. - Caveat: fragile command handling. Splitting on spaces breaks quoted arguments. Use a wrapper script for anything complex.
- Caveat: security defaults. A hard-coded gRPC key and unauthenticated file-sync routes mean it must not face the internet.
- Caveat: activity. The pinned commit is from October 2024, and the
0.6.0changelog entry still says “TBC”.
Sources: code at 0485310, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (20 pages), verified Q&A.
How it answers the AI web scraping questions
Each answer was drafted by a code-reading agent at commit 0485310. Its citations were checked mechanically. Compare with the other ai web scraping →
How are pages fetched and rendered?
answeredCrawlab does NOT fetch or render web pages itself. There is no HTTP client for scraping, no headless browser integration, no JavaScript rendering engine, and no fetching strategy in the codebase. The system is a management platform: it launches user-provided spider scripts as OS subprocesses, and those scripts do the actual page fetching via whatever framework the user chooses (Scrapy, Puppeteer, Selenium, raw HTTP, etc.). The RunnerV2 class (core/task/handler/runner_v2.go:88-169) reads the Cmd and Param fields from a Spider or Task model and runs the command via sys_exec.BuildCmd. The spider's Cmd field — e.g. scrapy crawl myspider or python main.py — is a freeform string the user sets in the Spider model (core/models/models/spider.go:28-31). The ConfigSpiderData entity defines a 'configurable spider' concept with stages and selectors, but the Scrapy code generator that would turn this into executable spider code is entirely commented out (core/models/config_spider/scrapy.go — every function body is commented out). No PDF, image, or other content-type handling is implemented in the platform. No content-type negotiation or JS-waiting strategy exists; those are entirely the concern of the user's spider script.
How is content extracted or converted?
answeredCrawlab does not extract or transform web content. There is no HTML to Markdown converter, no readability-style boilerplate removal, no CSS/XPath extraction engine, and no schema-based extraction runtime. Results flow from user spider scripts directly into a configurable data store. The ResultServiceMongo (core/result/service_mongo.go:42-82) inserts documents into a MongoDB collection specified by the spider's ColId/ColName. It optionally supports deduplication via hash-based duplicate checking on configurable keys (overwrite or skip modes). The ConfigSpiderData entity (core/entity/config_spider.go:3-41) defines a YAML/JSON schema with CSS/XPath selectors, stages, and field names — designed for a visual configurable spider feature, but the Scrapy code generator that would execute this schema is entirely commented out (core/models/config_spider/scrapy.go). The template-parser module (template-parser/general_parser.go:28-152) provides a {{variable}} template engine with JavaScript math evaluation using the Otto JS VM — a generic template renderer, not specific to web scraping. For results, Crawlab supports export to CSV and JSON (core/export/csv_service.go, core/export/json_service.go) by reading from the MongoDB data collection.
How are LLMs used, if at all?
not applicableCrawlab has no LLM integration of any kind. A case-insensitive search for llm, openai, chatgpt, gpt-, claude, langchain across all Go source files returned zero relevant results (only matches in protobuf-generated files and lock files unrelated to LLM use). There are no LLM providers configured, no prompt templates, no chunking logic for large pages, no structured output schemas backed by an LLM, and no cost-control mechanisms. The project does not use AI/ML for extraction, classification, summarization, or any other purpose. Scraping logic, data parsing, and field extraction are entirely the responsibility of user-written spider scripts running as subprocesses outside the Crawlab platform.
How are anti-bot measures, proxies and fingerprinting handled?
not applicableCrawlab implements no anti-bot countermeasures. There are no stealth patches, no browser fingerprint spoofing, no proxy rotation, no CAPTCHA handling, no rate-limiting of outgoing requests, no user-agent rotation, and no robots.txt parsing (grep for robots.txt across all Go files returns zero results). The DependencySetting.Proxy field (core/models/models/dependency_setting.go:16) is a proxy setting for downloading Python/Node package dependencies for spiders, not for scraping traffic — it is used only in the dependency-installation feature. The nginx configuration (nginx/crawlab.conf:14-17) reverse-proxies the API, which is unrelated to scraping. All anti-bot measures (proxies, browser automation stealth, cookie management, retry with backoff) must be implemented within the user's spider scripts, which execute outside Crawlab's control as opaque subprocesses.
How is crawling at scale implemented?
answeredCrawlab implements distributed task orchestration, not crawling itself, but it handles spider execution at scale. Architecture: Master and Worker nodes communicate via gRPC (core/grpc/server/task_server_v2.go:70-110). Task queue: Tasks are inserted as TaskV2 and TaskQueueItemV2 documents in MongoDB with a priority field (core/task/scheduler/service_v2.go:39-80). Workers fetch tasks via a gRPC Fetch RPC that queries and atomically dequeues from this queue, first by node assignment then by any-node fallback (core/grpc/server/task_server_v2.go:70-110). Concurrency: Each node has a configurable MaxRunners (default 8, core/node/config/config.go:20). The ServiceV2.Fetch() loop runs every second (core/task/handler/service_v2.go:76-130), checks available runner slots, and only fetches if capacity remains. URL dedup: Not implemented at the crawling level — deduplication only exists at the result-storage level (core/result/service_mongo.go:43-73, hash-based dedup on result documents). Politeness/robots.txt: None. Distributed workers: Run modes include all-nodes, random, and selected-nodes (core/constants/task.go:12-16). The master monitors workers every 15 seconds via gRPC pings (core/node/service/master_service_v2.go:93-114). File sync: Workers download spider files from master via HTTP before running (core/task/handler/runner_v2.go:329-436). Scheduling: A cron-based scheduler (core/schedule/service_v2.go:75-78) runs tasks on cron expressions. Cleanup: Tasks and stats older than 30 days are automatically purged (core/task/scheduler/service_v2.go:197-233).
What is the developer interface?
answeredCrawlab offers several interfaces. CLI: crawlab server (or crawlab s) starts the node service — master or worker determined by config (core/cmd/server.go:8-24). The root command has a -c flag for a custom config file (core/cmd/root.go:27). REST API: A Gin web server on port 8000 (core/apps/api_v2.go:46-73) exposes CRUD endpoints for spiders, tasks, schedules, projects, nodes, data collections, users, results, settings, and stats — all defined in core/controllers/router_v2.go:54-373. Login/auth via /login and /logout with JWT middleware (core/middlewares/auth.go). Web UI: A Vue.js frontend (frontend/src/main.ts) imports the crawlab-ui library (version 0.6.3) providing the full UI with Element Plus components, served by nginx on port 8080 (nginx/crawlab.conf:1-18). Export: Results can be exported as CSV or JSON (core/controllers/export_v2.go:12-27). gRPC: Internal protocol between master and worker nodes for task fetching, log streaming, and data transfer. Data sources: Results can be stored in MongoDB (default), MySQL, PostgreSQL, and others via the DatabaseV2 model (core/models/models/v2/database_v2.go:7-48) which supports Mongo, Postgres, Snowflake, Cassandra, Hive, and Redis connection parameters. Language bindings: None — spiders are arbitrary executables in any language, run as OS subprocesses.
edition is Pro, and the core/ds drivers are not imported anywhere at this SHA.