LLMs Technical Reviews
Home / AI web scraping / crawlab

crawlab-team/crawlab

Go and MongoDB platform that schedules and runs your own spider programs across nodes; it does no fetching or extraction itself.

GitHub ↗★ 12kGoBSD-3-Clausecommit 0485310 · 2024-10-09homepage ↗

Overview

Crawlab is not a scraper and has no AI features. It is a crawler management platform: a Go server, a MongoDB database and a web UI that run your spider programs on one or more machines. A spider can be a Scrapy project, a Node/Puppeteer script, a Go binary or any other command. Crawlab stores the code, starts it as an OS subprocess on a node, captures its logs and collects the records it reports back. Scheduling is done with cron expressions. Nothing in the repository fetches a web page, parses HTML or calls a model.

Where it fits in an AI scraping stack: it is the job-control layer. Your spider (perhaps a Crawl4AI or Firecrawl client, or a Playwright script that sends Markdown to an LLM) is the payload. Crawlab gives that payload a UI, cron schedules, per-node concurrency limits, live logs and a results table. Anything about fetching, stealth or extraction must live in the spider.

The pinned commit (October 2024) is the v0.6 rewrite. The code has *V2 services next to legacy ones, and some features are only half moved over. The sections below note where that matters.

Architecture

flowchart LR
  UI["Web UI (crawlab-ui package)"] -->|"/api via nginx"| API["Gin REST API"]
  API --> ADM["Spider admin service"]
  SCH["Cron schedule service"] --> ADM
  ADM --> Q["Mongo task queue"]
  API --> DB["MongoDB"]
  W["Worker task handler"] -->|"gRPC Fetch"| GS["Master gRPC server"]
  GS --> Q
  W --> R["Task runner"]
  R -->|"HTTP /sync"| API
  R --> P["Spider subprocess"]
  P -->|"gRPC stream: data, logs"| GS
  GS --> DB
  M["Master monitor"] -->|"ping every 15 s"| W
Component Path Role
Entry point core/cmd/server.go, core/apps/ crawlab server starts the API, gRPC server and node services
REST API core/controllers/router_v2.go Gin routes for spiders, tasks, schedules, nodes, results, export, sync
Spider admin core/spider/admin/service_v2.go Turns “run spider” into a task and a queue item
Task scheduler core/task/scheduler/service_v2.go Enqueue, cancel, status recovery, 30-day cleanup
Task handler core/task/handler/service_v2.go Per-node loop that fetches tasks and runs them
Task runner core/task/handler/runner_v2.go Syncs files, builds the command, sets env, starts and watches the process
gRPC task server core/grpc/server/task_server_v2.go Dequeues tasks; receives data and log streams
Master service core/node/service/master_service_v2.go Monitors workers, marks them offline
Result writer core/task/stats/service_v2.go Writes reported records to the spider’s collection
Schedules core/schedule/service_v2.go robfig/cron entries that call the spider admin
Frontend frontend/, nginx/crawlab.conf Loads the prebuilt crawlab-ui npm package; nginx serves it on 8080

How a request flows

Take a click on “Run” for a spider:

  1. API. PostSpiderRun binds SpiderRunOptions (mode, node ids, cmd, param, priority) and calls the spider admin service (spider_v2.go).
  2. Task creation. scheduleTasks builds one TaskV2 and fills empty fields from the spider’s defaults. If the mode named any nodes, it pins the task to the first one (service_v2.go). Enqueue inserts the task, a TaskQueueItemV2 with the priority and node id, and an empty stats row (scheduler/service_v2.go).
  3. Fetch. On every node, ServiceV2.Fetch ticks. It skips inactive or disabled nodes and nodes already at MaxRunners (default 8). Otherwise it calls the master’s gRPC Fetch (handler/service_v2.go). The master first looks for a queue item assigned to that node, then for an unassigned one. It sorts by priority and id, sets the task’s node and deletes the queue item (task_server_v2.go, L242-L263).
  4. Sync code. On a worker, RunnerV2.syncFiles asks the master’s /sync/:id/scan for a file list with hashes, deletes local files that are gone, and downloads changed files with 10 parallel requests (runner_v2.go).
  5. Build and start. configureCmd joins the command and param and passes them to sys_exec.BuildCmd. That function splits the string on single spaces, with no shell and no quoting (runner_v2.go, sys_exec_linux.go). configureEnv adds CRAWLAB_TASK_ID, the gRPC address and auth key, and every global environment variable from the database (runner_v2.go). Run starts the process, pipes stdout and stderr to the log stream, and waits on a signal channel for finish, cancel, error or “lost” (runner_v2.go).
  6. Report data. The spider uses the separate Crawlab SDK (save_item, or a Scrapy pipeline) to stream records over gRPC with the task id from the environment. Subscribe sends INSERT_DATA to handleInsertData and INSERT_LOGS to the log driver (task_server_v2.go, L212-L232). InsertData writes the batch into the spider’s Mongo collection and updates the result count (stats/service_v2.go).

Key components

Task queue

The queue is the task_queue_items Mongo collection, polled by every node through the master. This design is simple and needs no broker. Two details matter at scale. The dequeue is a find followed by a delete. It runs inside RunTransactionWithContext, but the model-service calls do not use the session context, so the transaction does not protect them. Two nodes polling at the same moment can claim the same item. And polling is the only dispatch: a task waits up to one fetch interval before a node picks it up.

Run modes

The UI offers “all nodes”, “random” and “selected nodes”. In the V2 admin service, all-nodes and selected-nodes both produce one task, assigned to the first node in the list. The legacy service created one task per node; that fan-out was not ported (legacy service.go). If you need the same job on every worker, schedule one task per node.

Node monitoring

The master runs monitor every 15 s. For each worker it subscribes, pings over gRPC and updates the free runner count. A worker that fails either check is marked offline (master_service_v2.go). The runner also has a process health check, but it returns at once when cmd.ProcessState is nil, and that is always the case while the process runs. So it never reports a lost process (runner_v2.go).

Results

Records go to the spider’s col_name collection in MongoDB. Only Pro editions (edition: global.edition.pro) route them to another database through a registry. The core/ds folder has MySQL, PostgreSQL, Elasticsearch, Kafka and other drivers, but no package imports it. Watch for one bug in the stats service. On a cache miss, getDatabaseServiceItem returns the nil item it looked up, not the entry it just cached. The first InsertData call for each task then dereferences nil (stats/service_v2.go). Results can be exported as CSV or JSON from the API.

Schedules

schedule/service_v2.go holds a robfig/cron instance. Each entry loads the schedule and the spider, merges their mode, nodes, cmd and param, and calls the same Schedule path as the Run button (schedule/service_v2.go).

Extending it

  • Any language. A spider is a folder plus a cmd string. Install the runtime on the node image. Global environment variables reach every task, which is a good place for API keys and proxy URLs.
  • Crawlab SDK. Spiders report records with save_item (Python) or CrawlabPipeline (Scrapy). Without the SDK, the platform still captures logs but sees no data.
  • Git-backed spiders. SpiderV2.GitId and GitRootPath sync code from a repository (vcs/ module).
  • REST API. Every UI action is a Gin route under /api, so CI or an agent can create spiders, start runs and read results.

Running it

  • Use Docker Compose: one crawlabteam/crawlab master (CRAWLAB_NODE_MASTER=Y), any number of workers pointing at it through CRAWLAB_GRPC_ADDRESS, and MongoDB. The UI is on port 8080. nginx forwards /api/ to the Go server on 8000. gRPC uses 9666.
  • The images are built from separate backend, frontend and plugin images. This repository’s frontend/ folder only bootstraps the crawlab-ui 0.6.3 npm package (main.ts).
  • Change the defaults. The gRPC auth key falls back to the constant Crawlab2021!. The /sync/:id/scan and /sync/:id/download routes are in the anonymous route group, and they join the path query onto the workspace directory without checking that it stays inside it (router_v2.go, sync_v2.go). Keep the API on a private network.

Strengths and caveats

  • Strength: framework-neutral. It runs anything with a command line, so it can schedule AI scrapers written in any stack without code changes.
  • Strength: operations UI. Live logs, task history, cron schedules, per-node runner caps and a data browser come ready-made.
  • Strength: few moving parts. Only MongoDB and the Crawlab binary. Workers sync code from the master automatically.
  • Caveat: no scraping features. It has no fetching, proxies, stealth, robots.txt, URL dedup, extraction or LLM support. Crawl-level concerns stay inside each spider.
  • Caveat: V2 migration gaps. Run modes do not fan out, the health check is dead code, the ds drivers are unused, and the stats cache has a nil bug. The two-version code makes it hard to tell what is live.
  • Caveat: fragile command handling. Splitting on spaces breaks quoted arguments. Use a wrapper script for anything complex.
  • Caveat: security defaults. A hard-coded gRPC key and unauthenticated file-sync routes mean it must not face the internet.
  • Caveat: activity. The pinned commit is from October 2024, and the 0.6.0 changelog entry still says “TBC”.

Sources: code at 0485310, deepwiki-open wiki (11 pages), OpenDeepWiki wiki (20 pages), verified Q&A.

How it answers the AI web scraping questions

Each answer was drafted by a code-reading agent at commit 0485310. Its citations were checked mechanically. Compare with the other ai web scraping →

How are pages fetched and rendered?

answered

Crawlab does NOT fetch or render web pages itself. There is no HTTP client for scraping, no headless browser integration, no JavaScript rendering engine, and no fetching strategy in the codebase. The system is a management platform: it launches user-provided spider scripts as OS subprocesses, and those scripts do the actual page fetching via whatever framework the user chooses (Scrapy, Puppeteer, Selenium, raw HTTP, etc.). The RunnerV2 class (core/task/handler/runner_v2.go:88-169) reads the Cmd and Param fields from a Spider or Task model and runs the command via sys_exec.BuildCmd. The spider's Cmd field — e.g. scrapy crawl myspider or python main.py — is a freeform string the user sets in the Spider model (core/models/models/spider.go:28-31). The ConfigSpiderData entity defines a 'configurable spider' concept with stages and selectors, but the Scrapy code generator that would turn this into executable spider code is entirely commented out (core/models/config_spider/scrapy.go — every function body is commented out). No PDF, image, or other content-type handling is implemented in the platform. No content-type negotiation or JS-waiting strategy exists; those are entirely the concern of the user's spider script.

How is content extracted or converted?

answered

Crawlab does not extract or transform web content. There is no HTML to Markdown converter, no readability-style boilerplate removal, no CSS/XPath extraction engine, and no schema-based extraction runtime. Results flow from user spider scripts directly into a configurable data store. The ResultServiceMongo (core/result/service_mongo.go:42-82) inserts documents into a MongoDB collection specified by the spider's ColId/ColName. It optionally supports deduplication via hash-based duplicate checking on configurable keys (overwrite or skip modes). The ConfigSpiderData entity (core/entity/config_spider.go:3-41) defines a YAML/JSON schema with CSS/XPath selectors, stages, and field names — designed for a visual configurable spider feature, but the Scrapy code generator that would execute this schema is entirely commented out (core/models/config_spider/scrapy.go). The template-parser module (template-parser/general_parser.go:28-152) provides a {{variable}} template engine with JavaScript math evaluation using the Otto JS VM — a generic template renderer, not specific to web scraping. For results, Crawlab supports export to CSV and JSON (core/export/csv_service.go, core/export/json_service.go) by reading from the MongoDB data collection.

How are LLMs used, if at all?

not applicable

Crawlab has no LLM integration of any kind. A case-insensitive search for llm, openai, chatgpt, gpt-, claude, langchain across all Go source files returned zero relevant results (only matches in protobuf-generated files and lock files unrelated to LLM use). There are no LLM providers configured, no prompt templates, no chunking logic for large pages, no structured output schemas backed by an LLM, and no cost-control mechanisms. The project does not use AI/ML for extraction, classification, summarization, or any other purpose. Scraping logic, data parsing, and field extraction are entirely the responsibility of user-written spider scripts running as subprocesses outside the Crawlab platform.

How are anti-bot measures, proxies and fingerprinting handled?

not applicable

Crawlab implements no anti-bot countermeasures. There are no stealth patches, no browser fingerprint spoofing, no proxy rotation, no CAPTCHA handling, no rate-limiting of outgoing requests, no user-agent rotation, and no robots.txt parsing (grep for robots.txt across all Go files returns zero results). The DependencySetting.Proxy field (core/models/models/dependency_setting.go:16) is a proxy setting for downloading Python/Node package dependencies for spiders, not for scraping traffic — it is used only in the dependency-installation feature. The nginx configuration (nginx/crawlab.conf:14-17) reverse-proxies the API, which is unrelated to scraping. All anti-bot measures (proxies, browser automation stealth, cookie management, retry with backoff) must be implemented within the user's spider scripts, which execute outside Crawlab's control as opaque subprocesses.

How is crawling at scale implemented?

answered

Crawlab implements distributed task orchestration, not crawling itself, but it handles spider execution at scale. Architecture: Master and Worker nodes communicate via gRPC (core/grpc/server/task_server_v2.go:70-110). Task queue: Tasks are inserted as TaskV2 and TaskQueueItemV2 documents in MongoDB with a priority field (core/task/scheduler/service_v2.go:39-80). Workers fetch tasks via a gRPC Fetch RPC that queries and atomically dequeues from this queue, first by node assignment then by any-node fallback (core/grpc/server/task_server_v2.go:70-110). Concurrency: Each node has a configurable MaxRunners (default 8, core/node/config/config.go:20). The ServiceV2.Fetch() loop runs every second (core/task/handler/service_v2.go:76-130), checks available runner slots, and only fetches if capacity remains. URL dedup: Not implemented at the crawling level — deduplication only exists at the result-storage level (core/result/service_mongo.go:43-73, hash-based dedup on result documents). Politeness/robots.txt: None. Distributed workers: Run modes include all-nodes, random, and selected-nodes (core/constants/task.go:12-16). The master monitors workers every 15 seconds via gRPC pings (core/node/service/master_service_v2.go:93-114). File sync: Workers download spider files from master via HTTP before running (core/task/handler/runner_v2.go:329-436). Scheduling: A cron-based scheduler (core/schedule/service_v2.go:75-78) runs tasks on cron expressions. Cleanup: Tasks and stats older than 30 days are automatically purged (core/task/scheduler/service_v2.go:197-233).

Editor's note. Correction: the dequeue is not atomic. The read and delete in getTaskQueueItemIdAndDequeue do not use the transaction's session context, so two concurrent Fetch calls can pick the same item. Also, in the V2 spider admin service the all-nodes and selected-nodes modes enqueue a single task pinned to the first node; only the legacy service fanned out one task per node.

What is the developer interface?

answered

Crawlab offers several interfaces. CLI: crawlab server (or crawlab s) starts the node service — master or worker determined by config (core/cmd/server.go:8-24). The root command has a -c flag for a custom config file (core/cmd/root.go:27). REST API: A Gin web server on port 8000 (core/apps/api_v2.go:46-73) exposes CRUD endpoints for spiders, tasks, schedules, projects, nodes, data collections, users, results, settings, and stats — all defined in core/controllers/router_v2.go:54-373. Login/auth via /login and /logout with JWT middleware (core/middlewares/auth.go). Web UI: A Vue.js frontend (frontend/src/main.ts) imports the crawlab-ui library (version 0.6.3) providing the full UI with Element Plus components, served by nginx on port 8080 (nginx/crawlab.conf:1-18). Export: Results can be exported as CSV or JSON (core/controllers/export_v2.go:12-27). gRPC: Internal protocol between master and worker nodes for task fetching, log streaming, and data transfer. Data sources: Results can be stored in MongoDB (default), MySQL, PostgreSQL, and others via the DatabaseV2 model (core/models/models/v2/database_v2.go:7-48) which supports Mongo, Postgres, Snowflake, Cassandra, Hive, and Redis connection parameters. Language bindings: None — spiders are arbitrary executables in any language, run as OS subprocesses.

Editor's note. Correction: in the V2 result path, non-Pro editions always write results to the spider's MongoDB collection (task/stats/service_v2.go InsertData); SQL and other targets are only used when edition is Pro, and the core/ds drivers are not imported anywhere at this SHA.