AI Blog

LlamaIndex vs Haystack vs RAGFlow vs R2R

All four do hybrid search, graphs and agentic retrieval, so the feature table decides nothing. Two things do: whether the framework runs inside your process or arrives as a second production system with its own database, users and on-call — and where the document-parsing boundary sits, because that is what decides whether your best-quality path is open source, a per-page bill, or an integration you own. Pick the posture; the features converged eighteen months ago.

By Agentic AI Wiki 14 min read

Every one of these four does hybrid search, knowledge graphs, multimodal ingestion and agentic retrieval, which means the feature table you are about to build will decide nothing. Two things will. The first is whether the framework runs inside your own process or arrives as a second production system with its own database, its own users and its own on-call rotation. The second is where the document-parsing boundary sits, because that decides whether your highest-quality path is an Apache-2.0 artefact you host, a per-page bill, or an integration you own outright. Choose the posture. The capabilities converged eighteen months ago; the postures never will.

At a glance

Four projects that show up on the same shortlist and are not the same kind of thing.

ProjectLicenceShapeWhat it actually is
LlamaIndex MIT Library, in your process An event-driven workflow framework whose commercial centre of gravity is document parsing
Haystack Apache-2.0 Library, in your process A component-and-pipeline orchestration framework with opinions about structure, not vendors
RAGFlow Apache-2.0 Server, beside your process A deployable RAG engine built around its own document-understanding models
R2R MIT Server, beside your process An agentic retrieval system exposed as a REST API with users, collections and auth built in
Feature matrix: process shape, open parsing, non-developer surface and exit cost A four-by-four matrix. Rows are LlamaIndex, Haystack, RAGFlow and R2R. Columns are whether retrieval runs inside your own process, whether high-quality document understanding is in the open artefact, whether a non-developer surface ships with it, and how much it costs to leave. LlamaIndex runs in-process, keeps its best parser behind a paid service, ships no end-user interface, and has a medium exit cost because workflow code is yours. Haystack runs in-process, treats parsing as a swappable component, ships no interface in the open project, and has the lowest exit cost because pipelines are serialisable components. RAGFlow runs as a separate stack, ships DeepDoc in the Apache-2.0 artefact, ships a full web interface, and has a high exit cost because knowledge bases live in its own database. R2R runs as a separate REST service, bundles a workable ingestion pipeline, ships a separate application, and has a medium exit cost because the surface is an API. Where each project leans hardest Runs inside your process Parsing in the open artefact Surface for non-developers Cost of leaving LlamaIndex In-process LlamaParse is paid None Medium Haystack In-process Swappable slot None in the OSS Low: components RAGFlow Separate stack DeepDoc, Apache-2.0 Full web UI High: KBs stay put R2R REST service Bundled pipeline Separate app Medium: API-shaped Leans hardest here Partial Not where it competes
Nothing in this matrix is a feature. Every column is a consequence you live with for years.

One number worth mentioning and then setting aside: RAGFlow's GitHub star count is comfortably the largest of the four, ahead of LlamaIndex, with Haystack third and R2R smallest. That ranking measures how many people wanted a RAG system they could docker compose up and show a colleague — which is a real and legitimate want, and which tells you nothing about which of these belongs in a service you are going to run for three years. Star-sorted shortlists systematically favour the projects with a UI.

The split nobody puts in the comparison table

A RAG library inside your process versus a RAG server beside it Two lanes. In the upper lane the framework is a library: your service calls it in-process, so ingestion, retrieval and generation happen inside your own deployment, your own traces and your own release cycle, with only the vector store and model provider outside. In the lower lane the framework is a server: your service calls it over HTTP, and behind that boundary sit the server's own ingestion pipeline, its own database, its own user and collection model and its own web interface, each with a release cycle you do not control. Where the retrieval failure shows up at 2am A library — LlamaIndex, Haystack Your deployment, your traces, your release Your service Calls the framework in-process. Ingest + retrieve Components you chose and wired. Generate Your prompts, your model choice. Vector store + model API The only things outside your process. A server — RAGFlow, R2R Your service Calls an HTTP API and waits. HTTP A second production system: its own deploy, upgrade cadence and on-call Ingestion Parsers bundled with the server. Its own database Documents, chunks, graphs, history. Users + collections A second identity model to reconcile. Its own UI Where non-devs do the work. Vector store + model API Configured inside the server, not in your code. The feature lists converge. This picture does not, and it is the part you cannot change later without a migration.
Same inputs, same outputs. The difference is who gets paged when retrieval returns the wrong chunk.

LlamaIndex and Haystack are libraries. You import them, wire components, and the retrieval runs in your process — same deploy, same logs, same trace, same language runtime, same dependency graph. When a query returns the wrong passage, the investigation happens in the code you already have open, and the fix ships with your next release.

RAGFlow and R2R are systems. You deploy them — RAGFlow as an orchestrated set of containers with a Python server and a TypeScript front end, R2R as a service you can start with pip install r2r for a lightweight run or via Docker Compose for the full thing — and your application talks to them over HTTP. Behind that boundary sits a second database holding documents, chunks, graphs and conversation history; a second identity model with its own users and collections; and, in RAGFlow's case, a full web interface that non-developers will start using directly. All of that is real capability. It is also a second production system, and you acquire its upgrade cadence, its backup story, its scaling behaviour and its incident surface along with its features.

The distinction matters most at the two moments teams do not plan for. The first is debugging: a library failure is a stack frame, a server failure is a support ticket against a component whose internals you did not write. The second is leaving. Migrating off a library means rewriting the wiring; migrating off a server means exporting documents, chunk boundaries, graph edges, collection membership and permission state from a schema designed for that server's convenience. That asymmetry is what the "cost of leaving" column in the matrix above is measuring, and it is the reason Haystack scores lowest on it — pipelines are serialisable component graphs, so what you built is largely a description you can carry.

None of this makes servers the wrong answer. If your actual requirement includes a knowledge-base interface that a compliance team maintains without opening a pull request, RAGFlow gives you in an afternoon what a library will not give you in a quarter. The mistake is arriving at that answer by way of a feature comparison instead of choosing it deliberately.

Where the parsing boundary sits

Three places the document-understanding boundary can sit Three columns describing where high-quality document parsing lives. In the open artefact: RAGFlow ships DeepDoc, its own layout, table-structure and OCR models, inside the Apache-2.0 distribution, so the best path is self-hosted. As a component you choose: Haystack and R2R treat converters as swappable, shipping a workable default and leaving the quality decision and its bill to you. As a paid service: LlamaIndex keeps the open library for orchestration and sells LlamaParse for parsing, so the highest-quality path leaves your network and is priced per page. Who owns the parser In the open artefact RAGFlow ships DeepDoc: layout, tables and OCR inside Apache-2.0. A component you pick Haystack and R2R ship a workable default and keep the slot swappable. A paid service LlamaIndex keeps the library open and sells LlamaParse separately. What it costs you What it costs you What it costs you GPU-shaped infrastructure and a heavier deployment, but nothing leaves. The quality decision, and the integration work, are both yours to make. A per-page bill and an egress path for every document you index.
"Supports PDF" is true of all four and settles nothing. This is the question underneath it.

In a real corpus, the single largest quality lever is not the retriever and not the reranker. It is whether the parser understood that the number in the third column of a merged-cell table belongs to the row header two rows up. Everything downstream — chunking, embedding, reranking, generation — inherits that decision, and no amount of hybrid search recovers a table that was flattened into prose. The four projects have made three different bets about who owns that step.

RAGFlow put it in the artefact. DeepDoc is its own visual document-understanding stack — OCR, table-structure recognition, layout recognition, with a vision-model pipeline option for images inside PDFs and DOCX files — and it ships inside the Apache-2.0 distribution. That is the unusual position of the four: the best-quality path is the self-hosted path, with no per-page meter and no document leaving your network. The bill arrives as infrastructure instead, because running document-understanding models is a different hardware conversation from running a Python service.

LlamaIndex went the other way, and the direction is visible in its own product story: the open library is the orchestration layer, and the commercial centre of gravity has moved to LlamaParse, sold as an agentic document-processing service. This is a coherent business and a genuinely strong parser. It is also a structural commitment: your highest-quality ingestion path leaves your network and is priced per page, which is a procurement conversation and a data-residency conversation before it is an engineering one. If either of those is hard at your organisation, discover it in week one rather than after the pipeline works.

Haystack and R2R both treat parsing as a slot. Haystack is explicit about it — converters are components like any other, swappable by design, and the framework declines to bundle a heavyweight model — which means you can point it at whichever parser wins next year, and also that reaching the quality ceiling is your integration project. R2R bundles a workable multimodal ingestion pipeline behind its API, which is less work up front and less visible when it is the thing costing you recall.

The practical consequence: run your own documents through all four ingestion paths before you compare anything else. Take twenty of your worst PDFs — the scanned ones, the ones with nested tables, the ones with two-column layouts and footnotes — and look at what comes out. That comparison takes a day and it discriminates between these projects far better than any query-quality benchmark will, because it is measuring the step whose errors nothing downstream can undo. The general problem is covered in document parsing for RAG.

Four projects saying "agentic", meaning four different things

All four have added agent language since 2024, and the word is now nearly content-free. What each actually ships:

  • LlamaIndex — Workflows. An event-driven step-composition model, now the recommended way to build anything non-trivial, with ReAct and function-calling agents built on top of it and a deployment path for running workflows as services. This is the most general of the four: it is a way to write agent control flow, in which retrieval is one thing you can do. If you want retrieval and orchestration from one library, this is the one that means it.
  • Haystack — agents as pipeline components. An agent is assembled from a chat generator, a tool invoker and tools, inside the same pipeline model as everything else, and Hayhooks deploys pipelines and agents as REST endpoints — or as MCP tools, which is the detail worth noticing: your retrieval pipeline becomes callable from an agent you built somewhere else entirely. Haystack's bet is that the pipeline is the primitive and the agent is a use of it.
  • RAGFlow — a visual workflow builder. Multi-step agents with persistent memory and tool calling, assembled in a UI, plus a browser component that lets an agent navigate web pages and chat-channel integrations for Feishu, Discord, Telegram and Line. This is agent-building for people who are not going to write the control flow in Python, and it is the only one of the four that seriously targets that audience.
  • R2R — a research agent behind an endpoint. Multi-step retrieval and synthesis exposed as an API call, with user-defined tooling so you can extend what the agent can reach. The agent is a feature of the retrieval service rather than a framework you program.

Read that list as four answers to "where does control flow live". In LlamaIndex it lives in your Python. In Haystack it lives in a pipeline definition. In RAGFlow it lives in a diagram in a web app. In R2R it lives inside the service. Any of those can be right; they are very hard to swap later, because the control flow is the part you accumulate the most of.

When to pick which

Your situationPick LlamaIndex or Haystack if…Pick RAGFlow if…Pick R2R if…
Retrieval inside an existing serviceYes — this is the case libraries are forYou are adding a system, not a featureOnly if you want the API boundary deliberately
Hard documents: scans, dense tablesBudget for a parser, open or paidStrongest default, and it is self-hostedBundled and adequate; test it first
Non-developers must curate the corpusYou are building that UI yourselfShips one; this is its best argumentSeparate app, developer-shaped
Nothing may leave the networkHaystack: no bundled service dependencyYes — best quality stays localYes, self-hosted
Agent orchestration beyond retrievalLlamaIndex Workflows, clearlyOnly if the visual builder fits your teamThe agent is theirs, not yours
You expect to migrate within two yearsHaystack — components travelPlan the export now, not laterAPI surface eases it; the data does not

FAQ

Is a RAG framework still worth adopting when models have long context?

Yes, for a reason unrelated to context length: retrieval is how you enforce permissions and provenance. Long context changes how much you can pass and does not change the need to prove which document an answer came from or to keep tenant A's files out of tenant B's prompt. See effective versus advertised long context.

Can I use a server-shaped framework and keep my own orchestration?

Yes, and it is often the sane compromise — call RAGFlow or R2R for retrieval and keep agent control flow in your own code. Be honest that you are then paying the operational cost of a second system for its ingestion quality and its interface, which is a defensible trade as long as it is the trade you meant to make.

Which one is fastest?

The question rarely survives contact with a profile. In-process retrieval avoids a network hop, which matters at low latency budgets; beyond that, latency in all four is dominated by the embedding call, the vector store and the generation, none of which the framework chooses for you. Measure with your corpus rather than trusting anyone's benchmark, including this page's framing.

Does the licence difference between MIT and Apache-2.0 matter here?

For most adopters, no — both are permissive and both are routinely accepted. Apache-2.0 includes an explicit patent grant, which some legal teams prefer; MIT does not, which some legal teams have never once raised. The licence question worth asking about a RAG project is not MIT-versus-Apache but which capabilities are in the open artefact at all, which is the parsing question above.

What should I evaluate first?

Ingestion, on your own worst documents, before you look at anything else. Retrieval quality, reranking and prompt design are all recoverable later; a parser that silently destroyed your tables sets a ceiling that no downstream component can lift.

Further reading

On this wiki:

Project sources: