Three of these four tools do not look at robots.txt unless you ask them to, and three of them send a spoofed browser User-Agent out of the box. The fifteen-year-old incumbent they are all replacing — Scrapy — is the only one in this comparison whose generated project obeys robots.txt, identifies itself with a URL, and caps concurrency at eight requests per domain. Pick on output quality and you will get a tool that works beautifully for a week and then stops, because the thing that ends a crawl is never extraction quality. It is the defaults you never read.
At a glance
Star counts are the values GitHub displayed on 10 October 2026; licences were read from the LICENSE files rather than taken from the sidebar label, which matters for two of the four.
| Project | Layer | Licence, from the file | Default output |
|---|---|---|---|
| Firecrawl — 190k | Managed HTTP API with a self-hostable core | AGPL-3.0 for the server; SDKs MIT | Markdown |
| Crawl4AI — 85.1k | Local Python library, plus a Docker server and a new cloud | Apache-2.0 plus a mandatory attribution rider | Markdown, with a filtered variant |
| ScrapeGraphAI — 31.7k | LLM-driven extraction graph on LangChain | MIT, unmodified | Schema-validated JSON |
| Spider — 2.8k | Low-level Rust crawling engine | MIT, uniform across all eight crates | Raw HTML; markdown via transform |
The defaults are the product
Read these from the source rather than the marketing. Crawl4AI's CrawlerRunConfig carries check_robots_txt: bool = False, with a docstring that states the default plainly; out of the box it does not fetch robots.txt at all. Spider's own unit test asserts its defaults — !config.respect_robots_txt, config.delay == 0, config.user_agent.is_none() — and its concurrency falls back to twenty or thirty permits with no per-origin cap, so the shipped configuration is tens of simultaneous requests at zero delay against a single host.
ScrapeGraphAI is the strangest case. It ships a RobotsNode, and no graph in the library wires it in, so every documented pipeline reaches a site without touching robots.txt. Wire it in yourself and you discover it is not a robots parser: it fetches the file, hands it to a language model, and asks whether it is "legit to scrape or not the website" — with an instruction to reply "yes" when the file is not provided. A non-deterministic compliance check that fails open is worse than none, because it produces a log line that looks like diligence.
Firecrawl is the only one of the four that tries, and it half-succeeds. Its /crawl endpoint defaults ignoreRobotsTxt to false, parses Crawl-delay, and identifies its robots token as FireCrawlAgent. Its single-URL /scrape path does not check robots at all unless a flag is set — which contradicts the README's blanket claim that "by default, Firecrawl respects robots.txt directives". Worth knowing too: overriding the robots user-agent is gated as an enterprise feature in the AGPL code.
Set against that, Scrapy's defaults are instructive precisely because everyone misquotes them. The framework default really is ROBOTSTXT_OBEY = False — but scrapy startproject writes ROBOTSTXT_OBEY = True into your settings file uncommented, ships a self-identifying User-Agent carrying a URL, and sets CONCURRENT_REQUESTS_PER_DOMAIN = 8. None of the four newer tools has a per-origin cap at all. That is the regression nobody advertises: the LLM era optimised hard for getting the page and dropped the operational etiquette that kept crawlers tolerated.
The licence label is wrong twice
Crawl4AI's GitHub sidebar says "Apache-2.0 license". The file is the Apache 2.0 text through END OF TERMS AND CONDITIONS, then a rule, then a section headed Attribution Requirement stating that all distributions, publications or public uses "must include" a specific credit line, "displayed in a prominent and easily accessible location" — a NOTICE or README for software, the acknowledgments for a paper, an About or Credits section for a website, the help output for a CLI. That is Apache-2.0 plus a further condition. It is not copyleft and not source-available, but it is not the bare permissive grant the label implies, and the README manages to contradict itself inside two paragraphs: one sentence says attribution "is recommended", the next subsection says you "must" include it. If your legal review turns on that word, get it resolved in writing rather than inferring it.
Firecrawl's label is accurate and incomplete in a different way. The server is AGPL-3.0, the language SDKs carry their own MIT files, and the README says so. What the licence cannot tell you is that the best scraping engine is not in the repository: fire-engine ships only as an HTTP client, reached through an environment variable, and the self-hosting guide states plainly that self-hosted scraping is "bundled Playwright with basic fetch fallback". So the self-hosted build is deliberately weaker than the hosted product on exactly the sites that are hard — which is the normal shape of an open-core business, and a thing to discover before the trial rather than after.
ScrapeGraphAI and Spider are the clean ones. Both are unmodified MIT with a single licence file; Spider symlinks the same file into all eight workspace crates and declares license = "MIT" in every manifest. Neither has a contributor licence agreement, and nor do the other two. If licence risk is your binding constraint, that is the whole answer and you can stop reading at this section.
Whether a model runs per page is the other fork
Three of the four are deterministic text pipelines: fetch, render if needed, convert to markdown, hand it to you. A model enters later, once per question you ask of the corpus. ScrapeGraphAI inverts that. An LLM is mandatory — AbstractGraph reads config["llm"] with an unguarded dict access, so there is no model-free code path — and the model runs on every page, turning your prompt into schema-validated JSON.
This is a real architectural choice rather than a flaw, and it changes three things at once. Cost moves from scaling with questions to scaling with pages, which is the difference between a bill you can bound and one that tracks your crawl frontier. Reproducibility goes: re-running the same crawl can disagree with the previous run, and because the pipeline's product is rows rather than text, the source page you would need to adjudicate the disagreement was never kept. And the schema is decided at crawl time, so changing your mind later means crawling again.
Pick the per-page shape when the corpus is small, heterogeneous and you genuinely cannot write selectors — a few thousand supplier pages with no two layouts alike. Pick the markdown shape when the corpus is large, or when you will ask it more than one question, or when anyone will ever need to see the page a fact came from. Crawl4AI deserves credit for making the cheap path the documented one: its CSS, XPath and regex extraction strategies run with no model at all, and the README's worked examples lead with them.
Two operational details that only matter once you have chosen. ScrapeGraphAI applies undetected-playwright stealth patches unconditionally on both async scrape paths, so it actively works to look like a human browser — a reasonable default for its use case and a hard conversation if your compliance team reads the dependency list. And its telemetry is on by default, with the flag defaulting to True; it is the only one of the four that phones home unless told otherwise, and the opt-out is an environment variable.
There is no usable performance comparison, and you should stop looking
Every number published in this category is vendor-run, and the vendors disagree with each other by margins that cannot both be true. Spider's own 1,000-URL test reports Spider at 99.9% success, Firecrawl at 95.3% and Crawl4AI at 89.7%. A benchmark run by a competing scraping API reports Firecrawl at 60.47% average success. Those two cannot be reconciled, and no independent replication of either exists.
Spider's headline throughput figures deserve the same scepticism even though its engineering is genuinely serious. Its own BENCHMARKS.md reports crawling a 185-page site in 73 ms on an M1 Max and 50 ms on a two-core CI runner, against tens of seconds for Go's Colly and Node's crawler on the same harness. Fifty milliseconds for 185 pages over real network IO is roughly 3,700 pages per second on a shared runner, which is not a plausible figure for network-bound work, and a six-hundred-fold gap to a mature Go crawler suggests the comparison implementations are not doing equivalent work. The project itself warns that CI results are flaky and that a dedicated machine is needed. Read it as evidence that Spider is fast, not as a number.
Firecrawl's claims — "covers 96% of the web" and "P95 latency of 3.4s across millions of pages" — are about the cloud product, and the latency figure ships with no methodology. ScrapeGraphAI publishes no performance numbers at all, which is the most honest position available here.
What to do instead takes an afternoon: take 200 URLs from the sites you actually need, run all the candidates against them with default settings, and score extraction completeness by hand on a sample. The result will be specific to your corpus, which is the only place these tools genuinely differ, and it will cost less than reading any of the benchmarks above.
The attack surface is the server, not the library
Crawl4AI published at least ten security advisories in 2026 — and that figure is a verified floor from the first page of its advisories list, not a total. The headline one is CVE-2026-57572, rated CVSS 10.0: unauthenticated remote code execution via Chromium launch-argument injection through browser_config.extra_args, with the advisory noting that the Docker API was unauthenticated by default, so a single request yielded arbitrary command execution. It is patched in 0.9.0, which the project describes as a secure-by-default release of the Docker server: authentication on, loopback bind, tokens reissued, and a request trust boundary that rejects a long list of previously-accepted fields over the network. A later batch fixed, among others, a blind SSRF in the robots.txt fetch itself, where the compliance check was the vulnerability.
Every one of those is in the self-hosted HTTP server, not the in-process library. If you pip install crawl4ai and call it from your own code, the exposure is small; if you expose the Docker API, you need to be on 0.9.4 or later and you should assume anything before 0.9.0 is compromised. Firecrawl's 2026 pair is similar in character: a critical command injection in change-tracking diff generation where page content reached a shell, patched in 2.11.31, and an arbitrary file read through JSON Schema $ref expansion in extraction, patched in 2.11.32.
ScrapeGraphAI and Spider have published none. That is worth reading carefully: it is evidence of no coordinated disclosure process in view, not evidence of security. Crawl4AI shipped ten advisories for an attack surface that ScrapeGraphAI largely shares — a Python library with a browser, an optional server, and user-supplied configuration. A project that publishes advisories is a project where somebody is looking.
When to pick which
| Situation | Pick | Because |
|---|---|---|
| You need pages as markdown today and AGPL is acceptable or you will use the hosted API | Firecrawl | The most complete product, the only one that checks robots on crawl, and the only one carrying a usage disclaimer |
| You want it in your own process, with no service and no copyleft | Crawl4AI | A library first; model-free extraction strategies are documented and fast. Set check_robots_txt=True and a real User-Agent on day one |
| Throughput is the binding constraint, or you are building your own crawler | Spider | HTTP-first with Chrome only when needed, clean MIT across every crate, no cloud dependency. Every politeness knob exists and every one is off |
| A small heterogeneous corpus where you cannot write selectors | ScrapeGraphAI | Per-page LLM extraction is the right tool for layouts you cannot predict. Budget per page, turn the telemetry off, and keep the source HTML yourself |
| Extraction quality on articles matters more than fetching | trafilatura, behind any of the above | The best boilerplate removal in the open-source world, and the only one with academic benchmarking — but no JavaScript rendering |
| You need crawl orchestration and a well-behaved client, not LLM-ready output | Scrapy | Per-domain concurrency, autothrottle, a self-identifying agent, fifteen years of operational knowledge |
Whichever you choose, the first commit after installing it is the same: turn robots checking on, set a User-Agent that names you and links to a page explaining what you are doing, and put a concurrency cap on each origin rather than on the fleet. None of the four gives you the third one, so it belongs in the fetch service you wrap around them. The enforcement that matters is moving to the network and CDN layer regardless of what any of these defaults say, and a crawler that cannot be identified is a crawler that cannot be allowlisted.
FAQ
Does the AGPL in Firecrawl affect me if I just call the hosted API?
No. Calling a hosted HTTP API is use, not distribution or modification, so the server's licence does not reach your code; the SDKs you link are MIT in their own licence files. The AGPL question only arises if you self-host and modify the server, or offer it to others over a network.
Is Crawl4AI's attribution rider a real obligation?
The LICENSE file says "must include" and specifies where the credit has to appear, so treat it as a condition rather than a courtesy. The ambiguity is that one README sentence calls attribution "recommended" while the file and the rest of the README say "must". A NOTICE line costs nothing; getting the ambiguity resolved in writing costs one email.
Which one respects robots.txt out of the box?
Only Firecrawl, and only on its crawl endpoint. Its single-page scrape path does not check by default, and Crawl4AI, ScrapeGraphAI and Spider all default to off. In Crawl4AI and Spider it is one configuration flag; in ScrapeGraphAI the node exists but no shipped graph uses it, and the node asks a language model rather than parsing the file.
Why is Spider's star count so much lower if it is this capable?
It is a Rust library at the engine layer, and the audience for that is a fraction of the audience for a Python one-liner or a hosted API. Popularity in this category tracks how little code it takes to get your first markdown page, which is close to the inverse of how much control the tool gives you.
Can I use more than one?
Routinely, and it is often the right answer: a fast engine for the frontier, a quality extractor for article bodies, a per-page model only on the pages that defeat your selectors. Keep the raw HTML in your own store so you can change the extraction layer later without re-crawling.
Should I worry that two of these have no published security advisories?
Treat it as an unknown, not a clean bill of health. All four expose a similar surface — a browser, user-supplied configuration, and in most cases an optional HTTP server — and the one that published ten advisories in 2026 is the one with somebody actively looking. If you expose any of these over a network, the version you are on matters more than the project you chose.
Further reading
On this wiki:
- Web-crawling and site-reading agents — why the deliverable is a fetch service rather than an instruction about politeness.
- Abuse reports about your agent — what happens when a site operator notices before you do.
- Bot verification and agent access — the identification layer that is replacing robots.txt as the thing that decides.
- Document parsing for RAG — the extraction-quality half of the problem, which no crawler solves for you.
- Egress control for agents — the proxy these tools should sit behind.
- Exa vs Tavily vs Brave Search vs Firecrawl — the same question one layer up, where the unit is a query rather than a page.
Project sources:
- firecrawl/firecrawl — note the owner has moved from
mendableai. - unclecode/crawl4ai — read
LICENSE, not the sidebar. - ScrapeGraphAI/Scrapegraph-ai
- spider-rs/spider — defaults asserted in
spider/src/configuration.rs. - scrapy/scrapy and adbar/trafilatura — the incumbents worth keeping in the comparison.