Every one of these four does hybrid search, knowledge graphs, multimodal ingestion and agentic retrieval, which means the feature table you are about to build will decide nothing. Two things will. The first is whether the framework runs inside your own process or arrives as a second production system with its own database, its own users and its own on-call rotation. The second is where the document-parsing boundary sits, because that decides whether your highest-quality path is an Apache-2.0 artefact you host, a per-page bill, or an integration you own outright. Choose the posture. The capabilities converged eighteen months ago; the postures never will.
At a glance
Four projects that show up on the same shortlist and are not the same kind of thing.
| Project | Licence | Shape | What it actually is |
|---|---|---|---|
| LlamaIndex | MIT | Library, in your process | An event-driven workflow framework whose commercial centre of gravity is document parsing |
| Haystack | Apache-2.0 | Library, in your process | A component-and-pipeline orchestration framework with opinions about structure, not vendors |
| RAGFlow | Apache-2.0 | Server, beside your process | A deployable RAG engine built around its own document-understanding models |
| R2R | MIT | Server, beside your process | An agentic retrieval system exposed as a REST API with users, collections and auth built in |
One number worth mentioning and then setting aside: RAGFlow's GitHub star count is comfortably the largest of the four, ahead of LlamaIndex, with Haystack third and R2R smallest. That ranking measures how many people wanted a RAG system they could docker compose up and show a colleague — which is a real and legitimate want, and which tells you nothing about which of these belongs in a service you are going to run for three years. Star-sorted shortlists systematically favour the projects with a UI.
The split nobody puts in the comparison table
LlamaIndex and Haystack are libraries. You import them, wire components, and the retrieval runs in your process — same deploy, same logs, same trace, same language runtime, same dependency graph. When a query returns the wrong passage, the investigation happens in the code you already have open, and the fix ships with your next release.
RAGFlow and R2R are systems. You deploy them — RAGFlow as an orchestrated set of containers with a Python server and a TypeScript front end, R2R as a service you can start with pip install r2r for a lightweight run or via Docker Compose for the full thing — and your application talks to them over HTTP. Behind that boundary sits a second database holding documents, chunks, graphs and conversation history; a second identity model with its own users and collections; and, in RAGFlow's case, a full web interface that non-developers will start using directly. All of that is real capability. It is also a second production system, and you acquire its upgrade cadence, its backup story, its scaling behaviour and its incident surface along with its features.
The distinction matters most at the two moments teams do not plan for. The first is debugging: a library failure is a stack frame, a server failure is a support ticket against a component whose internals you did not write. The second is leaving. Migrating off a library means rewriting the wiring; migrating off a server means exporting documents, chunk boundaries, graph edges, collection membership and permission state from a schema designed for that server's convenience. That asymmetry is what the "cost of leaving" column in the matrix above is measuring, and it is the reason Haystack scores lowest on it — pipelines are serialisable component graphs, so what you built is largely a description you can carry.
None of this makes servers the wrong answer. If your actual requirement includes a knowledge-base interface that a compliance team maintains without opening a pull request, RAGFlow gives you in an afternoon what a library will not give you in a quarter. The mistake is arriving at that answer by way of a feature comparison instead of choosing it deliberately.
Where the parsing boundary sits
In a real corpus, the single largest quality lever is not the retriever and not the reranker. It is whether the parser understood that the number in the third column of a merged-cell table belongs to the row header two rows up. Everything downstream — chunking, embedding, reranking, generation — inherits that decision, and no amount of hybrid search recovers a table that was flattened into prose. The four projects have made three different bets about who owns that step.
RAGFlow put it in the artefact. DeepDoc is its own visual document-understanding stack — OCR, table-structure recognition, layout recognition, with a vision-model pipeline option for images inside PDFs and DOCX files — and it ships inside the Apache-2.0 distribution. That is the unusual position of the four: the best-quality path is the self-hosted path, with no per-page meter and no document leaving your network. The bill arrives as infrastructure instead, because running document-understanding models is a different hardware conversation from running a Python service.
LlamaIndex went the other way, and the direction is visible in its own product story: the open library is the orchestration layer, and the commercial centre of gravity has moved to LlamaParse, sold as an agentic document-processing service. This is a coherent business and a genuinely strong parser. It is also a structural commitment: your highest-quality ingestion path leaves your network and is priced per page, which is a procurement conversation and a data-residency conversation before it is an engineering one. If either of those is hard at your organisation, discover it in week one rather than after the pipeline works.
Haystack and R2R both treat parsing as a slot. Haystack is explicit about it — converters are components like any other, swappable by design, and the framework declines to bundle a heavyweight model — which means you can point it at whichever parser wins next year, and also that reaching the quality ceiling is your integration project. R2R bundles a workable multimodal ingestion pipeline behind its API, which is less work up front and less visible when it is the thing costing you recall.
The practical consequence: run your own documents through all four ingestion paths before you compare anything else. Take twenty of your worst PDFs — the scanned ones, the ones with nested tables, the ones with two-column layouts and footnotes — and look at what comes out. That comparison takes a day and it discriminates between these projects far better than any query-quality benchmark will, because it is measuring the step whose errors nothing downstream can undo. The general problem is covered in document parsing for RAG.
Four projects saying "agentic", meaning four different things
All four have added agent language since 2024, and the word is now nearly content-free. What each actually ships:
- LlamaIndex — Workflows. An event-driven step-composition model, now the recommended way to build anything non-trivial, with ReAct and function-calling agents built on top of it and a deployment path for running workflows as services. This is the most general of the four: it is a way to write agent control flow, in which retrieval is one thing you can do. If you want retrieval and orchestration from one library, this is the one that means it.
- Haystack — agents as pipeline components. An agent is assembled from a chat generator, a tool invoker and tools, inside the same pipeline model as everything else, and Hayhooks deploys pipelines and agents as REST endpoints — or as MCP tools, which is the detail worth noticing: your retrieval pipeline becomes callable from an agent you built somewhere else entirely. Haystack's bet is that the pipeline is the primitive and the agent is a use of it.
- RAGFlow — a visual workflow builder. Multi-step agents with persistent memory and tool calling, assembled in a UI, plus a browser component that lets an agent navigate web pages and chat-channel integrations for Feishu, Discord, Telegram and Line. This is agent-building for people who are not going to write the control flow in Python, and it is the only one of the four that seriously targets that audience.
- R2R — a research agent behind an endpoint. Multi-step retrieval and synthesis exposed as an API call, with user-defined tooling so you can extend what the agent can reach. The agent is a feature of the retrieval service rather than a framework you program.
Read that list as four answers to "where does control flow live". In LlamaIndex it lives in your Python. In Haystack it lives in a pipeline definition. In RAGFlow it lives in a diagram in a web app. In R2R it lives inside the service. Any of those can be right; they are very hard to swap later, because the control flow is the part you accumulate the most of.
When to pick which
| Your situation | Pick LlamaIndex or Haystack if… | Pick RAGFlow if… | Pick R2R if… |
|---|---|---|---|
| Retrieval inside an existing service | Yes — this is the case libraries are for | You are adding a system, not a feature | Only if you want the API boundary deliberately |
| Hard documents: scans, dense tables | Budget for a parser, open or paid | Strongest default, and it is self-hosted | Bundled and adequate; test it first |
| Non-developers must curate the corpus | You are building that UI yourself | Ships one; this is its best argument | Separate app, developer-shaped |
| Nothing may leave the network | Haystack: no bundled service dependency | Yes — best quality stays local | Yes, self-hosted |
| Agent orchestration beyond retrieval | LlamaIndex Workflows, clearly | Only if the visual builder fits your team | The agent is theirs, not yours |
| You expect to migrate within two years | Haystack — components travel | Plan the export now, not later | API surface eases it; the data does not |
FAQ
Is a RAG framework still worth adopting when models have long context?
Yes, for a reason unrelated to context length: retrieval is how you enforce permissions and provenance. Long context changes how much you can pass and does not change the need to prove which document an answer came from or to keep tenant A's files out of tenant B's prompt. See effective versus advertised long context.
Can I use a server-shaped framework and keep my own orchestration?
Yes, and it is often the sane compromise — call RAGFlow or R2R for retrieval and keep agent control flow in your own code. Be honest that you are then paying the operational cost of a second system for its ingestion quality and its interface, which is a defensible trade as long as it is the trade you meant to make.
Which one is fastest?
The question rarely survives contact with a profile. In-process retrieval avoids a network hop, which matters at low latency budgets; beyond that, latency in all four is dominated by the embedding call, the vector store and the generation, none of which the framework chooses for you. Measure with your corpus rather than trusting anyone's benchmark, including this page's framing.
Does the licence difference between MIT and Apache-2.0 matter here?
For most adopters, no — both are permissive and both are routinely accepted. Apache-2.0 includes an explicit patent grant, which some legal teams prefer; MIT does not, which some legal teams have never once raised. The licence question worth asking about a RAG project is not MIT-versus-Apache but which capabilities are in the open artefact at all, which is the parsing question above.
What should I evaluate first?
Ingestion, on your own worst documents, before you look at anything else. Retrieval quality, reranking and prompt design are all recoverable later; a parser that silently destroyed your tables sets a ceiling that no downstream component can lift.
Further reading
On this wiki:
- Document Parsing for RAG — the step this whole comparison turns on.
- Advanced RAG Architectures — what these frameworks are implementations of.
- Choosing a Vector Database — the decision underneath all four.
- Evaluating RAG — how to make the bake-off mean something.
- Agentic Retrieval — what the word means once you take it seriously.