AI Blog

Firecrawl vs Crawl4AI vs ScrapeGraphAI vs Spider: read the defaults, not the README

Three of the four ignore robots.txt out of the box and three send a spoofed browser User-Agent; none has a per-origin concurrency cap, which the fifteen-year-old incumbent they are replacing has had all along. Two of the four also have a licence label that does not match the file. Pick on those, because every published benchmark here is vendor-run and they contradict each other.

By Agentic AI Wiki 16 min read

Three of these four tools do not look at robots.txt unless you ask them to, and three of them send a spoofed browser User-Agent out of the box. The fifteen-year-old incumbent they are all replacing — Scrapy — is the only one in this comparison whose generated project obeys robots.txt, identifies itself with a URL, and caps concurrency at eight requests per domain. Pick on output quality and you will get a tool that works beautifully for a week and then stops, because the thing that ends a crawl is never extraction quality. It is the defaults you never read.

At a glance

Star counts are the values GitHub displayed on 10 October 2026; licences were read from the LICENSE files rather than taken from the sidebar label, which matters for two of the four.

ProjectLayerLicence, from the fileDefault output
Firecrawl — 190k Managed HTTP API with a self-hostable core AGPL-3.0 for the server; SDKs MIT Markdown
Crawl4AI — 85.1k Local Python library, plus a Docker server and a new cloud Apache-2.0 plus a mandatory attribution rider Markdown, with a filtered variant
ScrapeGraphAI — 31.7k LLM-driven extraction graph on LangChain MIT, unmodified Schema-validated JSON
Spider — 2.8k Low-level Rust crawling engine MIT, uniform across all eight crates Raw HTML; markdown via transform
The sidebar label against the licence file Four columns comparing what GitHub's licence label says with what the LICENSE file in each repository actually contains. Firecrawl is AGPL for the server with MIT SDKs, Crawl4AI's Apache-2.0 carries a mandatory attribution rider, and ScrapeGraphAI and Spider are clean MIT. What the label says, and what the file says Firecrawl Label: AGPL-3.0 Server AGPL, SDKs MIT, best engine not in the repository at all Network copyleft plus open core Crawl4AI Label: Apache-2.0 Apache text plus a mandatory attribution rider the label omits Permissive, with a condition ScrapeGraphAI Label: MIT Unmodified MIT text, one licence file in the whole tree Label and file agree Spider Label: MIT Unmodified MIT, every crate symlinked to the same file Label and file agree Two of the four need the file read; two do not. The label alone decides nothing.
Two of these need the file opened before you can answer a procurement question. Two do not.

The defaults are the product

Out-of-the-box politeness defaults A matrix of five crawlers against four politeness defaults: whether robots.txt is obeyed, whether the crawler identifies itself, whether there is a per-origin concurrency cap, and whether requests to the target are rate-limited. Scrapy's project template is the only row with strong values; the four LLM-era tools default to off on nearly everything. What each tool does if you change nothing Obeys robots.txt Identifies itself Per-origin cap Throttles the target Firecrawl Crawl yes, scrape no No bot UA found No Opt-in delay Crawl4AI No Spoofed Chrome 116 No, global only No ScrapeGraphAI No, node unwired Stealth patches on No No Spider No Spoofed browser UA No, 20–30 global No, delay 0 Scrapy template Yes Yes, with a URL Yes, 8 per domain Opt-in autothrottle on by default partial or opt-in off by default The 15-year-old incumbent is the only row a site operator would recognise as a well-behaved client.
What each tool does if you install it and run the first example in its README.

Read these from the source rather than the marketing. Crawl4AI's CrawlerRunConfig carries check_robots_txt: bool = False, with a docstring that states the default plainly; out of the box it does not fetch robots.txt at all. Spider's own unit test asserts its defaults — !config.respect_robots_txt, config.delay == 0, config.user_agent.is_none() — and its concurrency falls back to twenty or thirty permits with no per-origin cap, so the shipped configuration is tens of simultaneous requests at zero delay against a single host.

ScrapeGraphAI is the strangest case. It ships a RobotsNode, and no graph in the library wires it in, so every documented pipeline reaches a site without touching robots.txt. Wire it in yourself and you discover it is not a robots parser: it fetches the file, hands it to a language model, and asks whether it is "legit to scrape or not the website" — with an instruction to reply "yes" when the file is not provided. A non-deterministic compliance check that fails open is worse than none, because it produces a log line that looks like diligence.

Firecrawl is the only one of the four that tries, and it half-succeeds. Its /crawl endpoint defaults ignoreRobotsTxt to false, parses Crawl-delay, and identifies its robots token as FireCrawlAgent. Its single-URL /scrape path does not check robots at all unless a flag is set — which contradicts the README's blanket claim that "by default, Firecrawl respects robots.txt directives". Worth knowing too: overriding the robots user-agent is gated as an enterprise feature in the AGPL code.

Set against that, Scrapy's defaults are instructive precisely because everyone misquotes them. The framework default really is ROBOTSTXT_OBEY = False — but scrapy startproject writes ROBOTSTXT_OBEY = True into your settings file uncommented, ships a self-identifying User-Agent carrying a URL, and sets CONCURRENT_REQUESTS_PER_DOMAIN = 8. None of the four newer tools has a per-origin cap at all. That is the regression nobody advertises: the LLM era optimised hard for getting the page and dropped the operational etiquette that kept crawlers tolerated.

The licence label is wrong twice

Crawl4AI's GitHub sidebar says "Apache-2.0 license". The file is the Apache 2.0 text through END OF TERMS AND CONDITIONS, then a rule, then a section headed Attribution Requirement stating that all distributions, publications or public uses "must include" a specific credit line, "displayed in a prominent and easily accessible location" — a NOTICE or README for software, the acknowledgments for a paper, an About or Credits section for a website, the help output for a CLI. That is Apache-2.0 plus a further condition. It is not copyleft and not source-available, but it is not the bare permissive grant the label implies, and the README manages to contradict itself inside two paragraphs: one sentence says attribution "is recommended", the next subsection says you "must" include it. If your legal review turns on that word, get it resolved in writing rather than inferring it.

Firecrawl's label is accurate and incomplete in a different way. The server is AGPL-3.0, the language SDKs carry their own MIT files, and the README says so. What the licence cannot tell you is that the best scraping engine is not in the repository: fire-engine ships only as an HTTP client, reached through an environment variable, and the self-hosting guide states plainly that self-hosted scraping is "bundled Playwright with basic fetch fallback". So the self-hosted build is deliberately weaker than the hosted product on exactly the sites that are hard — which is the normal shape of an open-core business, and a thing to discover before the trial rather than after.

ScrapeGraphAI and Spider are the clean ones. Both are unmodified MIT with a single licence file; Spider symlinks the same file into all eight workspace crates and declares license = "MIT" in every manifest. Neither has a contributor licence agreement, and nor do the other two. If licence risk is your binding constraint, that is the whole answer and you can stop reading at this section.

Whether a model runs per page is the other fork

Where the language model sits in a crawl Two pipelines compared. In the markdown path, the crawler fetches and converts pages deterministically and a model is called once per task downstream. In the per-page extraction path, a model is called for every page, which makes cost scale with pages crawled and makes the output non-reproducible. One model call per task, or one per page Markdown pipeline Firecrawl, Crawl4AI, Spider Fetch and render deterministic HTML to markdown deterministic Your store re-readable, diffable Model, once per question asked Cost scales with questions. Re-running the crawl gives byte-identical output. Per-page extraction pipeline ScrapeGraphAI Fetch and render deterministic Model, every page prompt in, JSON out Structured rows schema-validated Your store no page to re-read Cost scales with pages. A re-crawl can disagree with the last one, and the source text was not kept. The first shape lets you change your mind about the schema later. The second decides it at crawl time. Neither is wrong; they bill differently and fail differently.
The same crawl, billed two different ways and reproducible in only one of them.

Three of the four are deterministic text pipelines: fetch, render if needed, convert to markdown, hand it to you. A model enters later, once per question you ask of the corpus. ScrapeGraphAI inverts that. An LLM is mandatory — AbstractGraph reads config["llm"] with an unguarded dict access, so there is no model-free code path — and the model runs on every page, turning your prompt into schema-validated JSON.

This is a real architectural choice rather than a flaw, and it changes three things at once. Cost moves from scaling with questions to scaling with pages, which is the difference between a bill you can bound and one that tracks your crawl frontier. Reproducibility goes: re-running the same crawl can disagree with the previous run, and because the pipeline's product is rows rather than text, the source page you would need to adjudicate the disagreement was never kept. And the schema is decided at crawl time, so changing your mind later means crawling again.

Pick the per-page shape when the corpus is small, heterogeneous and you genuinely cannot write selectors — a few thousand supplier pages with no two layouts alike. Pick the markdown shape when the corpus is large, or when you will ask it more than one question, or when anyone will ever need to see the page a fact came from. Crawl4AI deserves credit for making the cheap path the documented one: its CSS, XPath and regex extraction strategies run with no model at all, and the README's worked examples lead with them.

Two operational details that only matter once you have chosen. ScrapeGraphAI applies undetected-playwright stealth patches unconditionally on both async scrape paths, so it actively works to look like a human browser — a reasonable default for its use case and a hard conversation if your compliance team reads the dependency list. And its telemetry is on by default, with the flag defaulting to True; it is the only one of the four that phones home unless told otherwise, and the opt-out is an environment variable.

There is no usable performance comparison, and you should stop looking

Every number published in this category is vendor-run, and the vendors disagree with each other by margins that cannot both be true. Spider's own 1,000-URL test reports Spider at 99.9% success, Firecrawl at 95.3% and Crawl4AI at 89.7%. A benchmark run by a competing scraping API reports Firecrawl at 60.47% average success. Those two cannot be reconciled, and no independent replication of either exists.

Spider's headline throughput figures deserve the same scepticism even though its engineering is genuinely serious. Its own BENCHMARKS.md reports crawling a 185-page site in 73 ms on an M1 Max and 50 ms on a two-core CI runner, against tens of seconds for Go's Colly and Node's crawler on the same harness. Fifty milliseconds for 185 pages over real network IO is roughly 3,700 pages per second on a shared runner, which is not a plausible figure for network-bound work, and a six-hundred-fold gap to a mature Go crawler suggests the comparison implementations are not doing equivalent work. The project itself warns that CI results are flaky and that a dedicated machine is needed. Read it as evidence that Spider is fast, not as a number.

Firecrawl's claims — "covers 96% of the web" and "P95 latency of 3.4s across millions of pages" — are about the cloud product, and the latency figure ships with no methodology. ScrapeGraphAI publishes no performance numbers at all, which is the most honest position available here.

What to do instead takes an afternoon: take 200 URLs from the sites you actually need, run all the candidates against them with default settings, and score extraction completeness by hand on a sample. The result will be specific to your corpus, which is the only place these tools genuinely differ, and it will cost less than reading any of the benchmarks above.

The attack surface is the server, not the library

Crawl4AI published at least ten security advisories in 2026 — and that figure is a verified floor from the first page of its advisories list, not a total. The headline one is CVE-2026-57572, rated CVSS 10.0: unauthenticated remote code execution via Chromium launch-argument injection through browser_config.extra_args, with the advisory noting that the Docker API was unauthenticated by default, so a single request yielded arbitrary command execution. It is patched in 0.9.0, which the project describes as a secure-by-default release of the Docker server: authentication on, loopback bind, tokens reissued, and a request trust boundary that rejects a long list of previously-accepted fields over the network. A later batch fixed, among others, a blind SSRF in the robots.txt fetch itself, where the compliance check was the vulnerability.

Every one of those is in the self-hosted HTTP server, not the in-process library. If you pip install crawl4ai and call it from your own code, the exposure is small; if you expose the Docker API, you need to be on 0.9.4 or later and you should assume anything before 0.9.0 is compromised. Firecrawl's 2026 pair is similar in character: a critical command injection in change-tracking diff generation where page content reached a shell, patched in 2.11.31, and an arbitrary file read through JSON Schema $ref expansion in extraction, patched in 2.11.32.

ScrapeGraphAI and Spider have published none. That is worth reading carefully: it is evidence of no coordinated disclosure process in view, not evidence of security. Crawl4AI shipped ten advisories for an attack surface that ScrapeGraphAI largely shares — a Python library with a browser, an optional server, and user-supplied configuration. A project that publishes advisories is a project where somebody is looking.

When to pick which

SituationPickBecause
You need pages as markdown today and AGPL is acceptable or you will use the hosted API Firecrawl The most complete product, the only one that checks robots on crawl, and the only one carrying a usage disclaimer
You want it in your own process, with no service and no copyleft Crawl4AI A library first; model-free extraction strategies are documented and fast. Set check_robots_txt=True and a real User-Agent on day one
Throughput is the binding constraint, or you are building your own crawler Spider HTTP-first with Chrome only when needed, clean MIT across every crate, no cloud dependency. Every politeness knob exists and every one is off
A small heterogeneous corpus where you cannot write selectors ScrapeGraphAI Per-page LLM extraction is the right tool for layouts you cannot predict. Budget per page, turn the telemetry off, and keep the source HTML yourself
Extraction quality on articles matters more than fetching trafilatura, behind any of the above The best boilerplate removal in the open-source world, and the only one with academic benchmarking — but no JavaScript rendering
You need crawl orchestration and a well-behaved client, not LLM-ready output Scrapy Per-domain concurrency, autothrottle, a self-identifying agent, fifteen years of operational knowledge

Whichever you choose, the first commit after installing it is the same: turn robots checking on, set a User-Agent that names you and links to a page explaining what you are doing, and put a concurrency cap on each origin rather than on the fleet. None of the four gives you the third one, so it belongs in the fetch service you wrap around them. The enforcement that matters is moving to the network and CDN layer regardless of what any of these defaults say, and a crawler that cannot be identified is a crawler that cannot be allowlisted.

FAQ

Does the AGPL in Firecrawl affect me if I just call the hosted API?

No. Calling a hosted HTTP API is use, not distribution or modification, so the server's licence does not reach your code; the SDKs you link are MIT in their own licence files. The AGPL question only arises if you self-host and modify the server, or offer it to others over a network.

Is Crawl4AI's attribution rider a real obligation?

The LICENSE file says "must include" and specifies where the credit has to appear, so treat it as a condition rather than a courtesy. The ambiguity is that one README sentence calls attribution "recommended" while the file and the rest of the README say "must". A NOTICE line costs nothing; getting the ambiguity resolved in writing costs one email.

Which one respects robots.txt out of the box?

Only Firecrawl, and only on its crawl endpoint. Its single-page scrape path does not check by default, and Crawl4AI, ScrapeGraphAI and Spider all default to off. In Crawl4AI and Spider it is one configuration flag; in ScrapeGraphAI the node exists but no shipped graph uses it, and the node asks a language model rather than parsing the file.

Why is Spider's star count so much lower if it is this capable?

It is a Rust library at the engine layer, and the audience for that is a fraction of the audience for a Python one-liner or a hosted API. Popularity in this category tracks how little code it takes to get your first markdown page, which is close to the inverse of how much control the tool gives you.

Can I use more than one?

Routinely, and it is often the right answer: a fast engine for the frontier, a quality extractor for article bodies, a per-page model only on the pages that defeat your selectors. Keep the raw HTML in your own store so you can change the extraction layer later without re-crawling.

Should I worry that two of these have no published security advisories?

Treat it as an unknown, not a clean bill of health. All four expose a similar surface — a browser, user-supplied configuration, and in most cases an optional HTTP server — and the one that published ten advisories in 2026 is the one with somebody actively looking. If you expose any of these over a network, the version you are on matters more than the project you chose.

Further reading

On this wiki:

Project sources: