Migrating off a memory provider.
A vector index is a pure function of your corpus and an embedding model: delete it, re-run the job, and you are back where you started. Agent memory is not, and that single difference is the whole of your exit risk. What a memory provider holds is a derived, path-dependent artefact — facts extracted by their prompts, deduplicated and superseded by their policy, in the order your history happened to arrive — and in the leading products that pipeline is explicitly the part that is not open. So the only asset you can actually carry out is the raw conversation history, which means the decision that determines whether you have an exit is made on the write path, before you integrate, and never again.
Memory is not an index, and the test is whether you can rebuild it.
Start by refusing the analogy that makes teams complacent. Re-indexing and embedding migrations is a painful, expensive operation, but it is a recomputation: every vector is determined by a chunk and a model version, both of which you hold, so the worst case is compute and wall-clock. Nothing is lost that cannot be regenerated.
A memory row is determined by at least five inputs: the raw history, the extraction prompt, the extraction model, the consolidation policy that decides what merges and what supersedes, and the prior memory state at the moment that episode arrived. You hold the first. The middle three belong to the provider. And the fifth makes the whole thing path-dependent — memory consolidates, contradicts, supersedes and forgets, which is the useful behaviour you are paying for and is also why replaying the same history in a different order produces a different store. There is no re-run that converges on what you have now.
So the practical question is not "can I export?" but "if I export, what have I got?" — and the honest answer is a snapshot of one provider's conclusions, with none of the reasoning and no way to reproduce it. That is a hedge against the provider disappearing. It is not a migration path.
A quick way to tell whether your team has understood this: ask what would happen if the provider silently shipped a better extraction model next week. For an index, nothing — your vectors are unchanged until you re-embed. For memory, your store begins accumulating facts under new rules while the old rows stay under the old ones, and nothing in your system records which is which. The absence of that boundary in your own data is the same absence that makes migration hard.
Read the licence for which capability is demo-grade, not for the SPDX identifier.
Nearly every serious project in this category ships under a permissive licence, and reading the headline tells you almost nothing about portability — because the boundary is drawn through the pipeline rather than through the code. Three concrete examples, all stated by the projects themselves as of October 2026:
- Mem0 is plain, unmodified Apache-2.0. Its own README nevertheless notes that the published benchmark figures come from the managed platform, which includes "proprietary optimizations not available in the open-source SDK", and that open-source users should expect "directionally similar gains but not identical numbers". The licence is genuinely open; the consolidation quality you measured in the trial may not be the thing you can self-host.
- Cognee is Apache-2.0, and ships a Postgres-as-graph-store path labelled a demo feature, with the note that "The production-ready version is available as a licensed product." Its bundled entity extractor carries the same caveat. Two of the capabilities you would build a production deployment on are deliberately demo-grade in the open repository.
- Zep publishes Graphiti, its temporal knowledge-graph framework, under Apache-2.0 while the product's own engine, managed deployment and user/thread management remain commercial. Self-hosting Graphiti is real and supported; it is bring-your-own graph database, and you build the surrounding operational tooling yourself.
None of this is sharp practice — it is the normal shape of an open-core business, and the projects document it in the open. The failure is on the buyer's side: a procurement checklist with a "licence" field gets a satisfying answer and stops, when the question that decides your exit is which specific capability is available at which tier. Put that question in the evaluation, in writing, before the trial: name the parts of the pipeline that differ between the edition we are testing and the edition we would run.
Exercise the export before you need it, and know what importing it does.
An export endpoint that has never been run at your volume is not an export endpoint. Run it in week one of the trial, at a realistic scale, and keep the output — it is both your hedge and your only honest documentation of what the provider considers yours.
- Check what the format carries. You will typically get extracted statements with timestamps, sometimes graph edges, sometimes derived user profiles. You will typically not get the extraction prompts, the supersession decisions, the embedding model version, or per-fact confidence. Without provenance, you cannot later tell a fact the agent was told from a fact the pipeline inferred.
- Time it and schedule it. An export that takes six hours and cannot be resumed is not usable in an incident. Find out during the trial, not during the migration.
- Understand that importing is re-extraction, not restoration. Loading one provider's conclusions into another's ingest path means its pipeline now treats somebody else's inferences as raw input. It extracts from them again, inheriting the first system's errors with no marker that they were inherited. Double-extracted memory is worse than cold memory, because it is confidently wrong in ways your evals were not designed to catch.
- Keep the export's retention in scope. A memory dump is a dense pile of personal data in a new location, which puts it squarely inside retention and legal hold, and makes it one more place a deletion request has to reach — see erasure against agent memory.
The one decision that preserves the exit: own the raw history.
Everything above argues for a single architectural rule, and it is cheap only if you adopt it before integrating. Write the raw transcript — and the tool results the memory was derived from — to your own durable store on the write path, independently of the provider call. Then the provider's store is a cache: expensive to warm, annoying to lose, and not load-bearing for your ability to leave.
Stated as a property: no memory provider should ever be the system of record for anything you could not reconstruct from your own data. If the only copy of a conversation lives in the provider, you do not have a vendor relationship, you have a dependency.
- Write yours first, or in parallel — never only after. A write path that stores locally only on success of the provider call loses exactly the history you will want when the provider is the problem.
- Store the inputs, not just the messages. Tool results are usually where the durable facts came from, they are frequently truncated before the model sees them, and they are the part nobody retains. Without them a replay reconstructs a different conversation.
- Do not reuse your trace store for this. Traces are sampled and short-retention by design — the correct design, per trace sampling and retention — which makes them the wrong substrate for the one copy that has to be complete and durable.
- Price it honestly, because it is the cheap half. Raw text is small and cold-storage priced; the extraction that turns it into memory is an LLM call per episode. Keeping the history costs storage. Not keeping it costs the option to leave.
- Decide the dual-write's failure mode on purpose. If your own store is unavailable, does the turn proceed? For the system of record the answer should be no — the fail-closed and fail-open question, with an unusually clear answer.
Run the cutover as a replay, and budget the bill it generates.
With the history in hand the migration stops being a data transfer and becomes a rebuild, which is slower, more expensive, and far more predictable. The sequence that works:
- Dual-write new episodes to both providers from your own store. Not from the old provider — from the source — so the new store is being built by its own pipeline from raw input, which is the only way its behaviour will match what you benchmarked.
- Replay history oldest-first, and budget it as a token bill. This is the step teams underestimate by an order of magnitude. Extraction is at least one model call per episode, so replaying a year of conversations for a large user base is a genuine spend with a genuine wall-clock, not an overnight copy. Rate-limit it so it does not contend with production traffic — the same discipline as any bulk and backfill run.
- Expect the new store to be worse on day one, by construction. It is empty, and it recovers as the replay lands. Teams read this as a quality verdict on the new provider and abort the migration three days in. Decide the measurement window before you start, and make it longer than the replay.
- Shadow-read on outcomes, not on overlap. You can compare what each provider retrieves per request, but there is no ground truth for "the right memories", so memory-overlap percentages will tell you only that the two systems differ. Score the downstream task instead, against recorded traffic — replay testing with recorded traces is the mechanism, and it is the only comparison that means anything here.
- Keep the old provider readable through the cutover and past it. Reads from both, writes to the new one, old one read-only — then keep that read path alive long enough to answer "what did it know in March", because somebody will ask.
Decide on your own eval, and never on the published scores.
This category's published numbers cannot arbitrate a buying decision, and the reasons are structural rather than a matter of any one vendor's honesty: different judge models, different reader models, different harnesses, and headline figures that are sometimes produced by a managed platform whose optimisations are not in the edition you would run. The same product has been reported at wildly different scores on a benchmark with the same name depending on who ran it and when — the memory benchmark is not the buying decision traces how that happens in detail. Treat any cross-vendor gap of a few points as noise, and any comparison of two vendors' self-reported figures as uninterpretable.
What replaces them is unglamorous and decisive: a frozen eval set drawn from your own recorded traffic, agreed before the migration starts, scored on task outcome. Fix the set, fix the judge, fix the reader, and then the only thing varying is the provider — which is the comparison you actually need and the one nobody can publish on your behalf.
- Agree the cutover threshold in advance. Mid-migration is the worst possible moment to negotiate what "good enough" means, and a replay that is still filling will always look like an argument for waiting.
- Separate retrieval failures from extraction failures. "The agent did not know" has two causes with opposite fixes, and the provider's dashboard will not distinguish them. Log the retrieved set alongside the answer so you can.
- Re-run the eval after the provider's next release. Extraction and consolidation change under you without a version you can pin, which is the unpinned vendor defaults problem with an unusually long feedback delay: the memory written this week is read next quarter.
- Report one number to your sponsor. Task success on the frozen set, old provider versus new, at the end of the agreed window. Everything else in this page is in service of making that one comparison honest.
If you are already integrated and none of this is in place, do the two cheap things this week and leave the rest. First, start writing raw transcripts and tool results to your own store — it is a write-path change, it costs storage, and from that day forward you have an exit you did not have yesterday. Second, run the provider's export at full volume once and time it, then diary it quarterly. The replay tooling, the frozen eval and the dual-write are the real work, but they are only worth building once the history exists; without it, there is nothing to replay and the migration you are planning is not possible at any budget.