AI Blog

GraphRAG vs LightRAG vs Graphiti vs Cognee: choose by write pattern

The retrieval quality gap between these four is far smaller than the gap in what an update costs, so the real decision is whether your graph is built once, appended to, or continuously mutated. And all four dedupe entities by string matching, which is the failure nobody's benchmark catches.

By Agentic AI Wiki 13 min read

The quality gap between these four graph-RAG systems is roughly the difference between a good answer and a slightly better one; the gap in what a single corpus update costs is two orders of magnitude. That asymmetry should decide the choice, and it means the question to answer first is not "which retrieves best" but "how often does my graph change, and what happens to it when it does". Then there is the part none of them solves: all four collapse duplicate entities by matching strings, and a missed duplicate is a false node your retrieval evaluation will never flag.

At a glance

Four projects that look like alternatives on a feature list and are actually answers to different questions. Sorted by the shape of the write, which is the axis that matters.

ProjectBuilt forWrite patternWhat an update costs
Microsoft GraphRAG Holistic questions over a static document corpus Batch build with community detection and summarisation Recomputing community hierarchies — effectively a re-index
LightRAG Graph-flavoured retrieval at ordinary RAG prices Incremental insert, no hierarchy above the graph Priced per new document, not per corpus
Graphiti Agent memory where facts stop being true Episodic mutation with edge invalidation A model call per episode, paid continuously
Cognee A memory engine you compose and constrain yourself Pipelines and DAGs over pluggable stores Whatever your pipeline re-runs — you decide the granularity
Where each system leans hardest, across four axes A matrix of four systems against four axes. Microsoft GraphRAG is weak on cheap incremental update, weak on handling time and contradiction, strong on corpus-wide theme questions, and medium on schema and ontology control. LightRAG is strong on incremental update, weak on time, medium on theme questions, and weak on schema control. Graphiti is strong on incremental update, strong on time and contradiction, weak on theme questions, and medium on schema control. Cognee is medium on incremental update, medium on time, medium on theme questions, and strong on schema and ontology control. No row is strong on every axis. Where each one leans hardest Cheap incremental update Time and contradiction Corpus-wide theme questions Schema and ontology control Microsoft GraphRAG Weak Weak Strong Medium LightRAG Strong Weak Medium Weak Graphiti Strong Strong Weak Medium Cognee Medium Medium Medium Strong Strong Medium Weak No row is strong on all four — the axes trade against each other by construction, not by maturity.
No row is strong on all four axes, and that is structural rather than a maturity gap — cheap updates and corpus-wide summaries pull against each other.

Microsoft GraphRAG — deep dive

What the expensive step buys

GraphRAG's distinguishing move happens at index time, not query time. After extracting entities and relations it runs community detection over the graph and then has a model write a summary of each community, and of communities of communities, producing a hierarchy of increasingly abstract descriptions of the corpus. That hierarchy is what lets it answer questions no chunk contains — "what are the recurring themes across these ten thousand incident reports" — because the answer is assembled from summaries rather than retrieved from text.

Why it is the wrong default

The same design makes updates brutal. Adding documents changes community membership, which changes the summaries, which changes the summaries above them; the honest operation is a rebuild, and published accounts of full indexing runs on large corpora reach four and five figures in model spend. Microsoft's own LazyGraphRAG variant exists precisely to defer that cost by pushing work to query time, which tells you how real the problem is. GraphRAG earns its keep on a corpus that holds still — a closed archive, a regulatory filing set, last quarter's tickets — and stops earning it the moment new documents arrive daily.

LightRAG — deep dive

Dual-level retrieval without the hierarchy

LightRAG keeps the graph and drops the community summarisation. Retrieval runs at two levels: low-level keys pull specific entities and their neighbourhoods, high-level keys pull broader themes, and the two sets are merged before generation. It is a smaller idea than GraphRAG's hierarchy and it retains most of the benefit for the queries people actually run, which are overwhelmingly local.

The cost argument, and its limit

Reported figures put LightRAG in the region of 70–90% of GraphRAG's answer quality at a small fraction of the indexing spend, with incremental inserts that touch only the new material. Treat the specific ratios as indicative — they depend heavily on corpus and model — but the direction is not in dispute, and it is large enough that the burden of proof sits with GraphRAG. The limit is the one it traded away: nothing in LightRAG revises an existing edge, and nothing summarises the corpus as a whole, so a genuinely global question degrades to a broader retrieval rather than a synthesis.

Graphiti — deep dive

Bi-temporal edges are the whole product

Graphiti is not a document-QA system wearing a graph. It ingests episodes — a conversation turn, a tool result, an event — and every edge carries a validity interval alongside the time it was ingested. When a new episode contradicts an existing fact, the old edge is invalidated rather than deleted, so the graph can answer both "what is true" and "what did we believe last March". For agent memory that is the requirement, because agent memory is a stream of assertions that go stale, not a corpus.

What it demands in return

Every write goes through a model call that decides what changed and what it invalidates, so ingestion cost scales with conversation volume rather than corpus size — a different bill shape that catches teams out. It also leans hard on reliable structured output; providers and smaller models without solid schema adherence produce malformed extractions and ingestion failures, which makes model choice a functional dependency rather than a tuning knob. See structured outputs.

Cognee — deep dive

A framework, not an opinion

Cognee is the most composable of the four: ingestion, normalisation and query are tasks and DAGs you assemble, and it plugs into several graph stores — Kuzu, Neo4j, FalkorDB — and several vector stores rather than shipping one. It also takes ontologies seriously, which matters more than it sounds: constraining extraction to a declared schema is the single most effective lever on graph quality, and it is the lever the other three mostly leave you to improvise.

The cost of composability

You inherit the decisions the other systems make for you. Chunking, extraction prompts, resolution strategy, what a re-run touches — all yours, which is exactly right if you have a domain schema worth enforcing and exactly wrong if you wanted a default that works by Friday.

Cross-cutting comparison

The write pattern is the decision

Three write patterns and what each one costs when the corpus changes Three columns. A batch build, used by Microsoft GraphRAG, re-derives communities and their summaries over the whole corpus, which gives the best corpus-wide answers but makes every update a re-index. An append, used by LightRAG and by Cognee's incremental pipelines, inserts new entities and edges without recomputing a hierarchy, which keeps updates cheap but weakens whole-corpus questions. A mutate, used by Graphiti, writes an episode and invalidates the edges it contradicts, which is the only pattern that models facts ceasing to be true, at the cost of a heavier per-write model call and a dependence on structured output. What happens when the corpus changes Build Microsoft GraphRAG re-derives communities and their summaries across the whole corpus. Buys: the best whole-corpus "what are the themes" answers. Costs: every update is a re-index. Assumes the corpus holds still. Append LightRAG inserts new entities and edges without recomputing any hierarchy above them. Buys: updates priced per document instead of per corpus. Costs: weaker global questions. Nothing revises an old edge. Mutate Graphiti writes an episode and invalidates the edges it contradicts, keeping both. Buys: the only pattern where a fact can stop being true. Costs: a model call per write. Leans on structured output.
Build, append, mutate. Every other difference between these systems follows from which one they chose.

Ask what your corpus does and the field sorts itself. A frozen archive queried for themes points at GraphRAG and nothing else does the job. A knowledge base that grows daily and is queried locally points at LightRAG, and paying GraphRAG's index cost repeatedly to get a marginally better answer is a bad trade you will notice on the invoice. A stream of user interactions where yesterday's preference is wrong today points at Graphiti, and neither of the first two will represent the contradiction at all — they will hold both facts and let the model pick. A domain with a real ontology and an engineering team willing to own the pipeline points at Cognee.

All four share the failure that matters

The stage all four systems share, and the one they all leave to you Four differently shaped inputs — a static corpus for Microsoft GraphRAG, an appended stream for LightRAG, timestamped episodes for Graphiti, and a composed pipeline for Cognee — all feed one LLM extraction step that emits entities and relations. That output passes through a deduplication step, highlighted as the critical stage, which in the popular systems is string matching and therefore misses abbreviations, synonyms, casing differences, translations and typos. The deduplicated triples land in a graph store, from which two retrieval paths run: local traversal from matched entities, and global answers over community or theme summaries. A band beneath notes that a missed duplicate creates a false node, which creates false edges, and that retrieval evaluations do not catch it because the answer stays fluent. Four write patterns, one shared weak point Static corpus Microsoft GraphRAG — batch build over everything Appended stream LightRAG — insert without a rebuild Timestamped episodes Graphiti — facts that later stop being true Composed pipeline Cognee — tasks and DAGs you define LLM extraction — entities and relations Prompt-driven, non-deterministic, priced per token of your whole corpus Deduplication — string matching Misses abbreviations, synonyms, casing, translations, typos. This is the stage every one of the four leaves to you. Graph store Kuzu, Neo4j, FalkorDB, or an embedded default Local traversal from matched entities outward Global answers over community or theme summaries A missed duplicate is not a missing row — it is a false node, which means false edges, which means the traversal walks a structure that does not match reality. The answer stays fluent, so a retrieval evaluation scored on answer quality will not tell you it happened.
The highlighted stage is the one every system leaves to you, and the one that decides whether the graph describes your domain or a slightly wrong version of it.

The popular graph-RAG implementations, LightRAG and GraphRAG among them, deduplicate extracted entities by string matching. That misses case differences, abbreviations, synonyms, translations and typos — so Acme Corp, Acme Corporation and ACME become three nodes. In a vector index a duplicate is a redundant row and costs you a little recall. In a graph it is a false node, which means false edges, which means every traversal through that region walks a structure that does not exist. The answer that comes back is still fluent and still cited, which is why answer-quality evaluations do not surface it; you find it by sampling the graph, not by scoring the output.

Budget for a resolution pass regardless of which system you pick: embedding-based candidate blocking, then a model or a rule set to confirm merges, run before the graph is queried rather than after someone complains. The research direction is well established — dedicated entity-resolution and triple-reflection stages measurably clean LLM-built graphs — but none of the four gives it to you turnkey today. Related: graph RAG for when the graph is worth building at all.

The graph store is not the choice

It is tempting to start from the database. Do not: Cognee is explicitly portable across Kuzu, Neo4j and FalkorDB, Graphiti runs on several backends, and the store rarely determines whether the system works. The extraction and resolution stages do. Our graph database comparison covers that layer separately, and it is a downstream decision.

When to pick which

SituationPickWhyWatch for
Closed archive, asked for themes and patterns Microsoft GraphRAG Community summaries are the only mechanism here that answers a whole-corpus question Budget the full index run before committing; re-indexing is the recurring cost
Knowledge base that grows daily, mostly local questions LightRAG Incremental inserts and a fraction of the indexing spend for most of the answer quality Global questions degrade to broad retrieval, not synthesis
Agent memory over conversations and events Graphiti Only pattern that models a fact ceasing to be true, via edge invalidation Ingestion cost scales with traffic; needs a model with reliable structured output
Regulated domain with a real ontology Cognee Schema-constrained extraction and pluggable stores; you own the pipeline You also own chunking, prompts and resolution — staff for it

And the option none of the four vendors will offer: build the graph only after a hybrid retrieval baseline has failed on a written-down query set. Graph construction adds an expensive, non-deterministic extraction stage to your pipeline, and a large share of questions people bring to graph RAG are answered by hybrid search with a reranker at a fraction of the operational cost.

FAQ

Is Microsoft GraphRAG obsolete?

No, but its niche is narrower than its mindshare. The community-summarisation hierarchy is still the strongest available mechanism for questions about a corpus as a whole, and nothing else here replicates it. It is the wrong default for a corpus that changes, because the update cost is a rebuild rather than an insert.

Can I use Graphiti for document question answering?

You can, and you will be paying for machinery you are not using. Graphiti's value is bi-temporal edges that let a fact be superseded — a property a static document set does not need. For documents, LightRAG or GraphRAG matches the workload better.

How much does entity resolution actually matter?

More than the choice between these four. String-match deduplication turns naming variants into separate nodes, and a false node produces false edges that every traversal through that region then follows. Because the generated answer stays fluent, output-quality evaluations do not detect it — you have to sample the graph itself.

Do I need a dedicated graph database?

Usually not to start. Cognee is portable across Kuzu, Neo4j and FalkorDB, Graphiti supports several backends, and embedded stores are fine at prototype scale. Choose the store after the extraction and resolution stages work, because those decide whether the system is useful.

Should I build a knowledge graph at all?

Only after a hybrid-search baseline has measurably failed on a written-down set of questions. Graph construction adds a non-deterministic, corpus-priced extraction stage, and many questions attributed to "we need a graph" are answered by better chunking and a reranker.

Further reading

On this wiki:

Project sources: