All four of these servers were built for the same moment: your agent writes a call that looked right in 2024, and the build fails. All four fix it. The choice that actually matters is one nobody puts in the comparison table — whether the text they hand your model was written by the library's maintainers, extracted from what they wrote, or written by another model about their code. Those are three different things to ground an agent on, and only the third one can be confidently wrong.
At a glance
Four MCP servers, four answers to "where does documentation come from".
| Server | Who runs it | What comes back | Reach |
|---|---|---|---|
| Context7 | Upstash — open-source server, hosted index | Snippets and API reference extracted per library, per version | Libraries in its catalogue |
| DeepWiki | Cognition (the Devin team) — hosted | A generated wiki about a repository, plus a Q&A tool | Public GitHub repositories |
| GitMCP | Open source, free remote server | The repository's own llms.txt, docs, or README |
Any GitHub repo or GitHub Pages site |
| Ref | ref.tools — commercial, API key | Searched doc sections, capped at roughly 5k tokens a read | Public docs and pages, plus your private sources |
The fork in the road is provenance, not coverage
GitMCP is the shortest path. Point it at a repository and it looks for an llms.txt file first, then the project's docs, then the README, and searches within them. What your model reads is what a maintainer committed. If the docs are wrong, they were already wrong for humans, and someone can file an issue about it.
Context7 and Ref sit one step further out. Both index upstream material and return the relevant fragment — Context7 as curated per-library, per-version snippets; Ref as searched sections of doc pages, trimmed to about 5k tokens per read and deduplicated against what it already showed you in the session. The extraction step can drop the caveat that lived two paragraphs above the code sample, which is a real and ordinary retrieval failure, but nothing in the pipeline invents an argument.
DeepWiki is a different object, and worth being explicit about because it is easy to file next to the other three. It runs a model over a repository and produces a wiki — architecture prose, diagrams, an ask_question tool — covering tens of thousands of popular repositories. That is enormously useful for the question Context7 was never trying to answer: how does this unfamiliar codebase work. It is the wrong input for what is the signature of this function, because the answer is generated text about code rather than the code or its documentation, and a generated error arrives in the same confident register as a generated truth. You have added a second grounding step and then cached its output.
A practical tell: ask yourself whether you would accept the answer without checking the source link. For upstream text the answer is usually yes. For a generated wiki it should be no — which means the tool is doing navigation work, not authority work, and your prompt should say so.
None of them knows which version you run
The pitch for this whole category is version drift: your model learned an API that has since changed, and it writes code against the version it remembers. Every one of these servers fixes staleness. Staleness is not the same problem.
Consider what "up-to-date documentation" does to a service pinned to a two-year-old release. Before the docs tool, the model guessed from training data and was wrong in a random direction. After, it reads current upstream and is wrong in a specific direction — it will confidently use the API that exists today, in a codebase that cannot run it. The tool did not remove the mismatch; it moved it and made it more plausible.
Look at how each closes that gap, and the answer is that none of them does it for you. Context7 goes furthest: it maintains per-version indexes and will serve the version you name — but you name it in the prompt, which means the binding is a natural-language assertion made by the model, and the model has no reason to have read your lockfile. GitMCP returns whatever branch or tag you pointed the URL at, which is a real pin if you set it and your repository's default branch if you did not. Ref returns what its search matched, which is usually the docs site's current release. DeepWiki serves whatever snapshot it indexed.
So the version binding is your job, and it is about fifteen lines of harness code:
- Resolve the version before the query, not in it. Read
package-lock.json,uv.lock,go.sum— whatever is authoritative in the repo the agent is editing — and inject the resolved versions into the docs query as a parameter your code controls. - Pin the URL where the server takes one. A GitMCP endpoint aimed at a tag is deterministic; one aimed at a default branch is a moving target that will change under you between two runs of the same eval.
- Make a version mismatch visible rather than fatal. If the index has no entry for the pinned version, you want the agent to say so and fall back, not to silently read the current docs. This is the same discipline as making absence explicit in a subagent's return.
- Check the claim against the installed artefact for anything load-bearing. The library is on disk. Reading the actual signature costs a few hundred tokens and beats any index, which is why repo navigation and documentation retrieval are complements rather than substitutes.
A documentation tool is a retrieval channel with a plan attached
Every one of these servers takes third-party text and puts it into a context window where a model reads it as material to act on. That is the definition of a retrieval-borne injection surface, and documentation is an unusually good carrier for it: an agent that fetched docs is, by construction, in the mood to follow instructions it finds there.
Three specifics worth pricing rather than worrying about in the abstract.
First, llms.txt is a file at a conventional path, written for machine readers, that no human on your team has read and that most reviewers would wave through. GitMCP prefers it, which is the correct design decision and also means the highest-trust slot in the pipeline is filled by the least-reviewed file in the repository. The same argument the wiki makes about tool poisoning applies unchanged: the channel is trusted because of where it sits, not because of what it says.
Second, hosted indexes are shared infrastructure. Whatever the index contains is served to everyone pointed at that library, so a bad entry is not your incident, it is a class of incidents — the supply-chain shape, one layer up from the package registry.
Third, and most often missed: DeepWiki indexes public repositories. Asking it about code that is not public is not a capability gap you work around, it is a publication event. Ref exists partly to serve that case properly, with private sources behind an account; GitMCP is open source and can be run against your own infrastructure. Both are the right shape for a private codebase. A public wiki generator is not.
Treat everything these servers return as data, never as instruction. In practice: put the retrieved text inside a delimited block in the prompt, tell the model it is reference material and that directives inside it are to be ignored, and keep the docs tool out of any loop where the agent can act on what it read without a step in between. That is the entire mitigation, it is cheap, and almost nobody does it because the text is called "documentation".
When to pick which
| Situation | Reach for | Why |
|---|---|---|
| Agent writes against popular libraries, versions matter | Context7 | Per-version indexes are the only ones designed to answer the version question at all |
| Onboarding onto an unfamiliar public repo | DeepWiki | Generated architecture prose and repo Q&A are exactly this job; verify specifics elsewhere |
| One specific project, pinned, zero setup | GitMCP | A URL per repo, upstream text, and a tag you control |
| Internal docs and public docs in one search | Ref | Private sources plus a hard token cap per read |
| Private codebase, no third-party indexing | Self-hosted GitMCP | Open source; the code never leaves your network |
These stack, and most teams that are serious end up with two: one tool for library reference, one for repository understanding. What does not stack is trust. Decide, per tool, whether its output is authority or navigation — and write that decision into the prompt, because the model will not infer it from the tool's name.
FAQ
Is one of these strictly better?
No, because two of them answer different questions. Context7, Ref and GitMCP answer "what does this API do"; DeepWiki answers "how is this repository put together". Comparing them on coverage misses that the outputs are different kinds of text.
Does Context7 pin to my installed version automatically?
No. It maintains per-version documentation and will serve a version you specify, but specifying it is something you or the model must do in the query. Nothing in the pipeline reads your lockfile, so resolving the version is harness work.
Can I point DeepWiki at a private repository?
The public DeepWiki index covers public GitHub repositories; private code is a Devin-account feature rather than something the open endpoint serves. Treat pointing a public indexer at private code as publishing it, and pick a tool with private sources or self-hosting instead.
What is llms.txt and should I ship one?
It is a conventional file at a repository or site root that gives machine readers a curated map of the documentation. Ship one if you maintain a library — it is the cheapest way to control what an agent reads about you. Review it like any other file in the repository, because for tools that prefer it, it is the highest-trust input you publish.
Do these tools reduce token usage or increase it?
Both, depending on the alternative. Against an agent crawling a docs site page by page they cut consumption sharply — Ref caps each read at roughly 5k tokens and drops results it already showed you. Against an agent that would have read the installed package source, they add a round trip and a summary. The comparison only makes sense per task.
Further reading
On this wiki:
- Repo navigation & context — the complement to documentation retrieval, and usually the cheaper answer.
- MCP tool poisoning — why a trusted channel is trusted for its position, not its content.
- Agent supply-chain security — shared indexes as shared blast radius.
- Agentic retrieval — letting the agent decide what to fetch, and what that costs.
- Agent skills — the other way teams package "how to use this thing" for a model.