The most popular project in this comparison no longer calls itself a research agent, and the second most popular has not taken a commit in over a year. LangChain archived open_deep_research on 21 August 2026 with 12.7k stars on it; DeerFlow's 2.0 rewrite re-described itself as a general agent harness; STORM's last commit on main is 30 September 2025. That is not four projects failing — it is a category dissolving, because plan-search-read-cite is now a default capability of any competent harness. What is left to choose on is narrower and more consequential: where your corpus lives, and who is allowed to see the query.
At a glance
Four projects that will all produce a cited report from a one-line prompt, sorted by what they are in October 2026 rather than by what they were when they got their stars.
| Project | Licence | Last moved | What it is now |
|---|---|---|---|
| GPT Researcher | Apache-2.0 | v3.7.0, 26 Sep 2026 | A research product: planner, executors, publisher, exports. |
| Local Deep Research | MIT | Commits on 6 Oct 2026 | A research app whose constraint is that nothing leaves the host. |
| STORM | MIT | Main commit 30 Sep 2025 | A frozen reference implementation of two good papers. |
| DeerFlow | MIT | v2.1.0, 24 Sep 2026 | A general agent harness; research is one thing it does. |
Read that chart as a history rather than a ranking. Stars accumulate from the moment a project is the clearest explanation of a new idea, and all four of these were, at different times, exactly that. None of them accrues stars for being the right dependency for your corpus today.
The loop stopped being the product
Open any of the four and you find the same pipeline: decompose the question into sub-questions, search, fetch and summarise each result, compress the accumulated findings so they fit a context window, then write a long answer with inline citations. LangChain's archived implementation is explicit about the middle stages even in its defaults — a cheap model for per-result summarisation, a stronger one for compression before the report is written.
Three independent signals say that pipeline is now infrastructure rather than a product:
- It got archived.
open_deep_researchwent read-only on 21 August 2026. LangChain's actively developed line isdeepagents, a batteries-included agent harness — the research loop moved down a layer into the thing that runs any agent. - It got absorbed. DeerFlow 2.0, released on 28 February 2026, is a ground-up rewrite that shares no code with v1, and it describes itself as a super agent harness that orchestrates sub-agents, memory and sandboxes "to do almost anything". The original research framework lives on the
1.xbranch. - It stopped needing changes. STORM's method was published at NAACL 2024 and EMNLP 2024, the repository added litellm support in January 2025, and main has been quiet since 30 September 2025. A frozen repository is not always abandonment; sometimes the contribution was the method.
This is the same consolidation that happened to retrieval-augmented generation itself, and it has the same consequence: the interesting engineering migrates to the ends of the pipeline. On the way in, which corpora you can reach and under what confidentiality. On the way out, whether a claim in the report is bound to the span that supports it — which, as citations and source-attribution UX sets out, is where generative search engines have historically been weakest.
GPT Researcher — the one that stayed a product
GPT Researcher (Apache-2.0, 29.9k stars, 4.1k forks) is the project that treated "deep research" as a thing to finish rather than a thing to generalise, and in a dissolving category that turns out to be the differentiated position.
What it actually gives you
A planner agent that generates research questions, execution agents that work them in parallel, and a publisher that assembles the report. The defaults target breadth — scrape 20+ sources, produce something past 2,000 words — and the output side is unusually finished for an open-source agent: Markdown, PDF and Word export, inline generated images, a deep-research mode that explores a tree of sub-questions rather than one fan-out, and an MCP server so another agent can call it as a tool.
The retriever surface is the real feature
Web search is the headline, but the list that matters is the local one: PDF, text, CSV, Excel, Markdown, PowerPoint and Word from a directory you point it at, alongside MCP-provided tools. That combination — your documents plus the web plus a tool protocol — is what makes it a plausible base for an internal research assistant instead of a demo.
Read its benchmark numbers as its own
v3.7.0 (26 September 2026) replaced embedding-based context filtering with a usefulness scorer the project calls Jev, and reports 73% relevant passages kept against 46% for embeddings over 28 research tasks. Treat that as a vendor-run A/B on a 28-task set — directionally interesting, not a field measurement — and note what it is measuring: the compression stage, which is the stage most likely to silently drop the paragraph your answer needed.
Local Deep Research — the one with a network boundary in the design
Local Deep Research (MIT, 9.2k stars, ~8,100 commits, commits landing on 6 October 2026) is the only one of the four whose central constraint is not quality or breadth but confidentiality: the queries, the documents and the model can all stay on hardware you own.
Local means local
Inference goes to Ollama, LM Studio or a llama-server endpoint from llama.cpp — three local backends, not a cloud SDK with a self-hosted escape hatch. Retrieval spans arXiv, PubMed and Semantic Scholar for academic work, SearXNG or SerpAPI for the open web, GitHub and Elasticsearch for technical corpora, plus local document collections, LangChain retrievers and the Wayback Machine. Per-user state sits in an encrypted database. With the cloud retrievers disabled, the whole thing runs inside the host, which is the property that makes it usable where a hosted research API is simply not an option — see air-gapped agent deployments for what else changes when you take that posture.
Its benchmark table is the most honest artefact in this comparison, and it undercuts the whole genre
The project publishes SimpleQA accuracy by local model on the same strategy: 95.7% with Qwen3.6-27B, 91.2% with Qwen3.5-9B, 85.4% with gpt-oss-20B, over 300–500 questions; and on xbench-DeepSearch (n=100), 77.0% against 59.0% for the same two Qwen models. Ten points of SimpleQA and eighteen points of DeepSearch moved without changing the project at all — only the weights behind it. Hold that number next to any cross-project comparison you read, including this one.
STORM — a method with a reference implementation, frozen
STORM (MIT, 31.6k stars) is the one to read rather than depend on. Its contribution is a pre-writing stage that most pipelines still skip: instead of fanning out sub-questions from the prompt, it discovers perspectives on the topic and then simulates conversations between a writer holding each perspective and an expert grounded in retrieved sources, using the resulting transcript to build an outline before a word of the article is drafted.
Why the outline-first order matters
Write-as-you-search produces a report whose structure is an artefact of search ordering, and that is visible in the output as repetition and as sections that exist because a query returned something. Outline-first forces the structure to be a decision, taken before retrieval has had a chance to bias it. Co-STORM extends this to a human-in-the-loop discourse protocol with a moderator and a maintained mind map, which is a better answer to "the user does not know what to ask" than a chat box.
Treat it as a paper with code
Main has been quiet since 30 September 2025 and the last release is v1.1.0 from January 2025, which added litellm so the model layer is at least pluggable. The papers (NAACL 2024, EMNLP 2024) and the FreshWiki and WildSeek datasets are the durable artefacts. Port the perspective-discovery and outline stages into whatever harness you are actually running; do not put a frozen repository in your dependency graph and call it a research platform.
DeerFlow — research as a skill inside a harness
DeerFlow (MIT, 83.4k stars, 11.6k forks) is where the category went. Version 2.0, released 28 February 2026, is a rewrite that shares no code with the research-only v1, and the positioning in the README is now a general harness: sub-agent delegation with per-agent scoping, LangGraph orchestration with checkpoint persistence, plan mode with TODO-status tracking and task attribution, manual context compaction, session goals, and MCP support covering stdio tools, HTTP/SSE servers and durable background tasks. Search providers are plugged in the same way everything else is.
What you get and what you take on
You get the thing the research-only projects do not have: a loop built for tasks that run for hours, with durable state and a plan you can inspect and interrupt. That matters, because a long research run is a durable-execution problem before it is a retrieval problem, and the research-shaped projects generally restart from scratch when the process dies.
You take on a harness. Sub-agents, memory, sandboxes and MCP servers are all surface you now own, operate and secure — and a research agent is a prompt-injection consumer by construction, since its entire job is to read untrusted web pages and act on them. Picking DeerFlow for research is picking a platform; that is the right call if you were going to need a platform anyway, and an expensive one if you needed a report generator.
Where the corpus lives decides everything else
Once the loop is commodity, the selection criterion is an exfiltration question with a quality question attached to it. Three regimes, and almost every real constraint sorts into one:
- Vendor index, vendor model. A hosted research API. Fastest to ship, best web coverage, and your query — which in an enterprise setting is often the sensitive part, not the answer — is a request to a third party. This is the research-API lane, and for public-web questions it is usually the correct one.
- Your corpus, their model. A self-hosted agent that calls a cloud model and a cloud search provider. Your documents never leave, but every chunk you retrieve is sent as a prompt to the model API and every sub-question is sent to a search provider. GPT Researcher and DeerFlow both sit here by default, and this is the regime most teams are actually in while believing they are in the third one.
- Your corpus, your model. Local Deep Research with the cloud retrievers off, or any of the others pointed at a local endpoint. Nothing leaves. You pay for it in model capability and in the operational burden of running inference, and the SimpleQA spread above tells you what the capability gap costs on easy questions.
The middle regime is the one worth being deliberate about, because it is a confidentiality decision that nobody makes explicitly — it arrives as a default base URL. If retrieved spans from an internal corpus are being sent to a model API, that is the fact to write in the design document, not discover in an audit; data governance and permission-aware retrieval are both about the gap between "the index is ours" and "the query is private".
What the feature matrix does and does not decide
Two things this matrix is good for: ruling out a project that cannot meet a hard constraint, and telling you how much platform you are adopting. One thing it is useless for: predicting whether the report will be any good. On that, the evidence in this comparison points at the model and the retriever, not at the orchestration code — a ten-point SimpleQA swing from changing local weights, inside one project, on one strategy.
So the efficient order of operations is: choose the confidentiality regime, pick the cheapest project that satisfies it, then spend your effort on the retriever list and the model, which is where the quality is. Reverse that order and you will spend a fortnight migrating orchestration code for a result that a different model would have given you in an afternoon.
When to pick which
| If this is your binding constraint | Start with | Why |
|---|---|---|
| Nothing may leave the host | Local Deep Research | Three local inference backends, local collections, encrypted per-user state; designed for this rather than configured into it. |
| An internal research assistant over your own documents | GPT Researcher | Local file retrievers plus web plus MCP, finished export paths, and it is still shipping as a research product. |
| You already run a harness, or will need one | DeerFlow (or your existing harness) | Durable state, sub-agent scoping, plan inspection. Research is a skill here, not the product. |
| Report structure is the thing that is wrong | STORM's method | Perspective discovery and outline-before-drafting, ported into your own loop. Read the papers; do not depend on the repository. |
| Public-web questions, no confidentiality constraint | A hosted research API | Coverage and latency you will not match, and no index to maintain. |
One cross-cutting rule, whichever row you land on: adopt the corpus connector, not the report generator. The connector is the part that encodes your permissions, your freshness and your confidentiality posture, and it is the part you cannot buy back later.
FAQ
Is an archived repository a reason not to use it?
It is a reason not to depend on it, which is different. Archived means read-only: no security fixes, no API-change follow-ups, and every model or SDK deprecation becomes your patch. Read the code, port the parts you need, and keep the dependency out of your lockfile.
Which of these produces the best report?
Unanswerable as asked, and the project's own numbers say why: within Local Deep Research, swapping local weights moved SimpleQA from 85.4% to 95.7% on an unchanged pipeline. The orchestration differences between these four are smaller than the differences between the models you put behind them.
Do I still need a vector database for this?
Often not. These loops are search-driven rather than index-driven — fetch, summarise, compress — and several leading agents dropped their index in favour of plain search over the corpus. If your corpus is large, private and queried repeatedly, an index starts paying; see local-first retrieval for where that line sits.
Why is DeerFlow so much more popular than the rest?
Because it stopped competing in this category. Stars follow general-purpose agent harnesses in 2026, which is the same market signal that archived one project here and froze another. It is not evidence that DeerFlow writes better research reports.
What about citation faithfulness — do these bind claims to sources?
All four emit citations; none of them guarantees that the cited source supports the sentence. That gap is the single highest-value thing to add yourself, and the measurement to run is planted-error detection rather than citation count — the argument is in research and synthesis agents.
Further reading
On this wiki:
- Research & Synthesis Agents — retrieve, write, verify, and citation faithfulness as a hard constraint.
- Local-First Retrieval — whether you need an index at all, and what to run if you do.
- Late-Interaction Retrieval — when the ranking quality you are missing is worth a bigger index.
- Permission-Aware Retrieval — what "our index" has to mean before a research agent can touch it.
- OpenAI vs Gemini vs Perplexity vs Exa research APIs — the hosted lane, for when the questions are public.
Project sources:
- assafelovic/gpt-researcher — Apache-2.0, 29.9k stars, v3.7.0 of 26 September 2026, the Jev context-filter figures and the retriever list.
- LearningCircuit/local-deep-research — MIT, 9.2k stars, the local-backend list and the SimpleQA and xbench-DeepSearch tables by model.
- stanford-oval/storm — MIT, 31.6k stars, v1.1.0 of January 2025, last main commit 30 September 2025, plus the NAACL 2024 and EMNLP 2024 papers.
- bytedance/deer-flow — MIT, 83.4k stars, the 2.0 rewrite of 28 February 2026, v2.1.0 of 24 September 2026, and the super-agent-harness positioning.
- langchain-ai/open_deep_research — MIT, 12.7k stars, archived read-only on 21 August 2026; the summarisation and compression defaults.