The quality gap between these four graph-RAG systems is roughly the difference between a good answer and a slightly better one; the gap in what a single corpus update costs is two orders of magnitude. That asymmetry should decide the choice, and it means the question to answer first is not "which retrieves best" but "how often does my graph change, and what happens to it when it does". Then there is the part none of them solves: all four collapse duplicate entities by matching strings, and a missed duplicate is a false node your retrieval evaluation will never flag.
At a glance
Four projects that look like alternatives on a feature list and are actually answers to different questions. Sorted by the shape of the write, which is the axis that matters.
| Project | Built for | Write pattern | What an update costs |
|---|---|---|---|
| Microsoft GraphRAG | Holistic questions over a static document corpus | Batch build with community detection and summarisation | Recomputing community hierarchies — effectively a re-index |
| LightRAG | Graph-flavoured retrieval at ordinary RAG prices | Incremental insert, no hierarchy above the graph | Priced per new document, not per corpus |
| Graphiti | Agent memory where facts stop being true | Episodic mutation with edge invalidation | A model call per episode, paid continuously |
| Cognee | A memory engine you compose and constrain yourself | Pipelines and DAGs over pluggable stores | Whatever your pipeline re-runs — you decide the granularity |
Microsoft GraphRAG — deep dive
What the expensive step buys
GraphRAG's distinguishing move happens at index time, not query time. After extracting entities and relations it runs community detection over the graph and then has a model write a summary of each community, and of communities of communities, producing a hierarchy of increasingly abstract descriptions of the corpus. That hierarchy is what lets it answer questions no chunk contains — "what are the recurring themes across these ten thousand incident reports" — because the answer is assembled from summaries rather than retrieved from text.
Why it is the wrong default
The same design makes updates brutal. Adding documents changes community membership, which changes the summaries, which changes the summaries above them; the honest operation is a rebuild, and published accounts of full indexing runs on large corpora reach four and five figures in model spend. Microsoft's own LazyGraphRAG variant exists precisely to defer that cost by pushing work to query time, which tells you how real the problem is. GraphRAG earns its keep on a corpus that holds still — a closed archive, a regulatory filing set, last quarter's tickets — and stops earning it the moment new documents arrive daily.
LightRAG — deep dive
Dual-level retrieval without the hierarchy
LightRAG keeps the graph and drops the community summarisation. Retrieval runs at two levels: low-level keys pull specific entities and their neighbourhoods, high-level keys pull broader themes, and the two sets are merged before generation. It is a smaller idea than GraphRAG's hierarchy and it retains most of the benefit for the queries people actually run, which are overwhelmingly local.
The cost argument, and its limit
Reported figures put LightRAG in the region of 70–90% of GraphRAG's answer quality at a small fraction of the indexing spend, with incremental inserts that touch only the new material. Treat the specific ratios as indicative — they depend heavily on corpus and model — but the direction is not in dispute, and it is large enough that the burden of proof sits with GraphRAG. The limit is the one it traded away: nothing in LightRAG revises an existing edge, and nothing summarises the corpus as a whole, so a genuinely global question degrades to a broader retrieval rather than a synthesis.
Graphiti — deep dive
Bi-temporal edges are the whole product
Graphiti is not a document-QA system wearing a graph. It ingests episodes — a conversation turn, a tool result, an event — and every edge carries a validity interval alongside the time it was ingested. When a new episode contradicts an existing fact, the old edge is invalidated rather than deleted, so the graph can answer both "what is true" and "what did we believe last March". For agent memory that is the requirement, because agent memory is a stream of assertions that go stale, not a corpus.
What it demands in return
Every write goes through a model call that decides what changed and what it invalidates, so ingestion cost scales with conversation volume rather than corpus size — a different bill shape that catches teams out. It also leans hard on reliable structured output; providers and smaller models without solid schema adherence produce malformed extractions and ingestion failures, which makes model choice a functional dependency rather than a tuning knob. See structured outputs.
Cognee — deep dive
A framework, not an opinion
Cognee is the most composable of the four: ingestion, normalisation and query are tasks and DAGs you assemble, and it plugs into several graph stores — Kuzu, Neo4j, FalkorDB — and several vector stores rather than shipping one. It also takes ontologies seriously, which matters more than it sounds: constraining extraction to a declared schema is the single most effective lever on graph quality, and it is the lever the other three mostly leave you to improvise.
The cost of composability
You inherit the decisions the other systems make for you. Chunking, extraction prompts, resolution strategy, what a re-run touches — all yours, which is exactly right if you have a domain schema worth enforcing and exactly wrong if you wanted a default that works by Friday.
Cross-cutting comparison
The write pattern is the decision
Ask what your corpus does and the field sorts itself. A frozen archive queried for themes points at GraphRAG and nothing else does the job. A knowledge base that grows daily and is queried locally points at LightRAG, and paying GraphRAG's index cost repeatedly to get a marginally better answer is a bad trade you will notice on the invoice. A stream of user interactions where yesterday's preference is wrong today points at Graphiti, and neither of the first two will represent the contradiction at all — they will hold both facts and let the model pick. A domain with a real ontology and an engineering team willing to own the pipeline points at Cognee.
All four share the failure that matters
The popular graph-RAG implementations, LightRAG and GraphRAG among them, deduplicate extracted entities by string matching. That misses case differences, abbreviations, synonyms, translations and typos — so Acme Corp, Acme Corporation and ACME become three nodes. In a vector index a duplicate is a redundant row and costs you a little recall. In a graph it is a false node, which means false edges, which means every traversal through that region walks a structure that does not exist. The answer that comes back is still fluent and still cited, which is why answer-quality evaluations do not surface it; you find it by sampling the graph, not by scoring the output.
Budget for a resolution pass regardless of which system you pick: embedding-based candidate blocking, then a model or a rule set to confirm merges, run before the graph is queried rather than after someone complains. The research direction is well established — dedicated entity-resolution and triple-reflection stages measurably clean LLM-built graphs — but none of the four gives it to you turnkey today. Related: graph RAG for when the graph is worth building at all.
The graph store is not the choice
It is tempting to start from the database. Do not: Cognee is explicitly portable across Kuzu, Neo4j and FalkorDB, Graphiti runs on several backends, and the store rarely determines whether the system works. The extraction and resolution stages do. Our graph database comparison covers that layer separately, and it is a downstream decision.
When to pick which
| Situation | Pick | Why | Watch for |
|---|---|---|---|
| Closed archive, asked for themes and patterns | Microsoft GraphRAG | Community summaries are the only mechanism here that answers a whole-corpus question | Budget the full index run before committing; re-indexing is the recurring cost |
| Knowledge base that grows daily, mostly local questions | LightRAG | Incremental inserts and a fraction of the indexing spend for most of the answer quality | Global questions degrade to broad retrieval, not synthesis |
| Agent memory over conversations and events | Graphiti | Only pattern that models a fact ceasing to be true, via edge invalidation | Ingestion cost scales with traffic; needs a model with reliable structured output |
| Regulated domain with a real ontology | Cognee | Schema-constrained extraction and pluggable stores; you own the pipeline | You also own chunking, prompts and resolution — staff for it |
And the option none of the four vendors will offer: build the graph only after a hybrid retrieval baseline has failed on a written-down query set. Graph construction adds an expensive, non-deterministic extraction stage to your pipeline, and a large share of questions people bring to graph RAG are answered by hybrid search with a reranker at a fraction of the operational cost.
FAQ
Is Microsoft GraphRAG obsolete?
No, but its niche is narrower than its mindshare. The community-summarisation hierarchy is still the strongest available mechanism for questions about a corpus as a whole, and nothing else here replicates it. It is the wrong default for a corpus that changes, because the update cost is a rebuild rather than an insert.
Can I use Graphiti for document question answering?
You can, and you will be paying for machinery you are not using. Graphiti's value is bi-temporal edges that let a fact be superseded — a property a static document set does not need. For documents, LightRAG or GraphRAG matches the workload better.
How much does entity resolution actually matter?
More than the choice between these four. String-match deduplication turns naming variants into separate nodes, and a false node produces false edges that every traversal through that region then follows. Because the generated answer stays fluent, output-quality evaluations do not detect it — you have to sample the graph itself.
Do I need a dedicated graph database?
Usually not to start. Cognee is portable across Kuzu, Neo4j and FalkorDB, Graphiti supports several backends, and embedded stores are fine at prototype scale. Choose the store after the extraction and resolution stages work, because those decide whether the system is useful.
Should I build a knowledge graph at all?
Only after a hybrid-search baseline has measurably failed on a written-down set of questions. Graph construction adds a non-deterministic, corpus-priced extraction stage, and many questions attributed to "we need a graph" are answered by better chunking and a reranker.
Further reading
On this wiki:
- Graph RAG — when a graph earns the extraction cost, and when it does not.
- Knowledge graphs — the concept, in plain terms.
- Memory stores — where a temporal graph sits among the alternatives.
- Evaluating RAG — why answer-quality scores hide structural errors like a false node.
- Neo4j vs Memgraph vs FalkorDB vs LadybugDB — the storage layer underneath all of this.