If your research output is going into a record rather than in front of a person, only one of these four can tell you which source supports which field — and that single difference outranks every comparison of report quality you will read. A search API returns documents and leaves the agent loop in your process. A research API takes the loop, which is the product and the price at the same time: you stop paying to orchestrate, and you stop being able to instrument.
At a glance
Four services that accept a question and return findings, with the loop on their side of the wire.
| Service | Interface | Default corpus | Billing unit |
|---|---|---|---|
OpenAI deep research (o3-deep-research, o4-mini-deep-research) | Responses API, background mode | None by default — you must attach web_search_preview and/or an MCP server | Tokens: $10/M in, $40/M out |
| Gemini Deep Research agent | Interactions API, async only, paid tier only | Google Search, URL Context, code execution | Per task: roughly $2 standard, $5 Max |
| Perplexity Sonar Deep Research | Chat-completions shaped | Perplexity's own index | Five meters, listed below |
| Exa Research / Exa Agent | REST, sync or async | Exa's neural index | Per request: roughly $12–15 per 1,000 |
Prices are list prices at the time of writing and all four have moved in the last six months; treat the ratios as durable and the absolute figures as perishable.
What you are actually buying
Build a research agent on a search API and you write the loop: decompose the question, issue queries, fetch and read, decide whether the evidence is sufficient, issue more queries, synthesise. It is perhaps three hundred lines, it is tedious, and it is where all your cost and all your failure modes live. Every one of those steps is also a span in your own trace, scored by your own evals.
A research API deletes those three hundred lines. In exchange, the plan, the query expansion, the read/skip decision and the stopping rule all move behind an API boundary you cannot see through. What comes back is a report and a list of sources. You can evaluate that report end to end — and end to end is the only way you can evaluate it, because the trajectory that produced it is not yours.
That is a legitimate trade and frequently the right one. It stops being right the moment someone asks why the answer is wrong, because the honest reply is that the report is a summary written by the system you are trying to debug. Everything below is about which of the four gives you the most to work with when that happens.
The axis nobody tables: what a citation is
All four advertise citations. They mean two different things by the word, and the gap between them is the single largest difference in this comparison.
A bibliography is not a claim-level guarantee
OpenAI's deep research models and Perplexity's Sonar Deep Research both return prose with a set of URLs attached. The URLs are real and were really visited. What they are not is a binding between a particular assertion and the evidence for it. If the report says a threshold is 40 milliseconds and lists nine sources, you cannot programmatically determine which of the nine — if any — contains that number. A human can check it by reading nine pages. A downstream system cannot check it at all.
Gemini sits in the middle. The Deep Research agent provides granular sourcing for claims rather than only a document list, which is materially better for a reader and still not a machine contract: you get an attribution, not a schema you asked for with a value in it.
Field-level grounding is a different product
Exa is the outlier and it is the reason this comparison exists. You pass an outputSchema; you get back output.content matching that schema and output.grounding carrying citations and a confidence per field. That is not a nicer citation format. It is the difference between a document you read and a record you can write into a database with a provenance column already populated — the thing that the preceding post on agent-supplied values in authoritative records argues you should never ship without.
The decision rule is short. If a human reads the output and acts on it, a bibliography is fine and you should optimise for report quality. If the output populates fields — a CRM record, a pricing table, a due-diligence sheet, an eval set — you need grounding bound to the field, and today that narrows the field to one. Retrofitting field-level grounding onto a prose report by asking a second model to align claims to sources is a known-bad pattern: you have added a second ungrounded generation step and called it verification.
Whose corpus, and whether you can extend it
The second real split is whether the loop can be pointed at anything of yours. It is a cleaner divide than the marketing suggests.
- OpenAI requires you to choose one. A deep research request must attach at least one data source:
web_search_preview, a remote MCP server, or both, andcode_interpreteris available alongside. The MCP server is not an ordinary one — the deep research models expect a specific shape, asearchtool returning top-k identifiers and afetchtool resolving those identifiers to documents. That constraint is doing real work: it forces your corpus into the same retrieve-then-read interface the web tool uses, which is why the loop can treat both sources identically. - Gemini supports remote MCP servers and file search, but not custom function calling. You can restrict the agent to the web (Google Search plus URL Context), to private sources (file search plus your MCP servers), or to a mix. What you cannot do is hand it an arbitrary function — so a research task that needs to call your pricing service mid-loop is out of scope, and the workaround is to wrap the service as an MCP server with search/fetch semantics it does not naturally have.
- Perplexity and Exa are the open web, full stop. Both are excellent at it and neither pretends otherwise. If your research question needs your wiki, your data room or your ticket history, these are the wrong tier of product and no amount of prompt context substitutes — you are asking the model to reason about documents it cannot go and get.
So the build-versus-buy question is not about quality. It is whether your research corpus is public. If it is, buy; the vendors are better at web research than your loop will be. If the answer lives half inside your company, you are choosing between OpenAI and Gemini on MCP ergonomics, or you are back to writing the loop over an agentic retrieval layer of your own.
The money: four incommensurable meters
Written out, the list prices at the time of writing are:
OpenAI o3-deep-research $10 / M input $40 / M output
Gemini deep-research ~$2 / task (standard), ~$5 / task (Max)
underneath: Gemini 3.1 Pro at $2 / M in, $12 / M out
plus Google Search grounding at $14 / 1K queries
(~80 queries standard, ~160 Max; implicit caching
covers 50-70% of input tokens)
Perplexity sonar-deep-research $2 / M input $8 / M output
+ $2 / M citation tokens
+ $3 / M reasoning tokens
+ $5 / 1K search queries
Exa research w/ schema ~$12-15 / 1K requests
agent fixed modes $0.012 (Minimal) - $1.00 (X-high) / run
These cannot be compared without fixing loop depth, and loop depth is the vendor's decision, not yours. A token-metered service has no ceiling: a question that triggers forty searches costs four times a question that triggers ten, and nothing in your request said how many to run. That is the whole difference between OpenAI's meter and Exa's.
- Gemini's per-task price is the most honest artefact in the group, precisely because the line items are published underneath it. Eighty grounding queries at $14 per thousand is $1.12 of search inside a $2 task, which tells you immediately that the inference is nearly free and the retrieval is the product. It also tells you the fixed price holds only while the query budget does.
- Perplexity's five meters are the most unpredictable, and the two unusual ones are the reason: citation tokens and reasoning tokens are volumes you neither specify nor see in advance. The prices are individually low. The variance is not.
- Exa's fixed modes buy predictability with a capped effort level. Minimal to X-high is a dial you set per run, which is the right shape for a production pipeline and the wrong shape for a question that deserves to run long because it turned out to be hard.
- Nobody bills for the retry. When a run returns a report that fails your quality gate, you pay again. Fold your gate's pass rate into every figure above before comparing anything — a service at 70% of the price with a 60% pass rate is more expensive, and this is the same arithmetic as cost per successful task everywhere else.
They are all jobs, not calls
Every one of these runs for minutes, and two of them refuse to pretend otherwise: Gemini's Deep Research agent is asynchronous only, and OpenAI's deep research models are meant to be driven in background mode because a synchronous request will not survive the run. Exa offers both shapes and the async one is the one you want above the cheapest modes.
The operational consequence is the part teams discover late. You need somewhere to persist the job handle, a poller or a webhook, a retry that does not re-run a completed job, a place to put a partial result when the run fails at minute nine, and a user-facing story for the wait. That is the durable-state infrastructure you were hoping to skip by buying the loop — buying the loop removes the orchestration of research steps, not the orchestration of the job.
Two smaller things that bite:
- The corpus is untrusted input and the loop reads it autonomously. A research agent's entire job is to fetch and read pages written by strangers, and three of these four will also follow links. Whatever comes back is data, not instruction, and the fact that the vendor owns the loop does not make it their problem — the report lands in your context window. See prompt injection.
- Reports are long, and length is a downstream cost. A 6,000-word report is 8,000 tokens in every subsequent turn of whatever agent consumes it. Ask for the schema you need rather than the essay, where the service lets you.
When to pick which
| Situation | Pick | Why |
|---|---|---|
| Output populates structured fields in a record | Exa | The only field-level grounding with a confidence; everything else needs a second ungrounded pass to fake it |
| The answer lives partly in internal systems | OpenAI or Gemini | Both take remote MCP servers; OpenAI forces a search/fetch shape, Gemini forbids custom functions |
| A human reads the report and decides | Gemini or OpenAI | Longest, best-synthesised output; per-claim sourcing on Gemini is the better reading experience |
| High volume, fixed budget per item | Exa or Gemini | Per-request and per-task pricing are the only two that let you forecast; token meters cannot be capped |
| Freshness matters more than depth | Perplexity | Purpose-built research API over a live index, and the fastest of the four to first result |
| You need the trajectory, not just the answer | None of them | Write the loop over a search API — this is the case the category cannot serve |
FAQ
Is a research API just a search API with a prompt on top?
No, and the difference is the loop rather than the prompt. A search API answers one query; a research API plans a sequence of queries, decides what to read, decides when it has enough, and writes the synthesis. You are buying the stopping rule as much as the search.
Which one produces the best report?
For long, hard, ambiguous questions read by a person, OpenAI and Gemini are consistently the deepest, and Perplexity is the fastest to something adequate. That ranking is also the least durable thing in this post — it has changed twice this year, and it is the wrong axis to architect on.
Can I get field-level citations out of the other three?
Not natively. The common workaround is a second model pass that maps report sentences to the source list, which produces confident alignments that are themselves ungrounded. If you need the guarantee, get it from the service that computes it during retrieval, not from a post-hoc aligner.
Do any of them read my private data?
OpenAI and Gemini can, through remote MCP servers — OpenAI requires a search/fetch interface specifically, and Gemini additionally offers file search while not supporting custom function calling. Perplexity and Exa search their own web indexes only.
How should I evaluate one of these?
Build a set of questions whose answers you already know and that are not answerable from a single page, then score the answer, the recall of sources you know exist, and the rate at which the report asserts something no cited source contains. Only the third is unusual and it is the one that separates the four. See why agent eval is hard.
Is it cheaper to write my own loop?
On a per-task basis, usually yes, once you already have a search API contract and a harness. On a total basis, usually no, until volume is high enough that your engineering time amortises — and the crossover is not about cost per call, it is about whether you need the trajectory for eval or debugging. That need, not the price, is the reason to keep the loop.
Further reading
On this wiki:
- Agentic retrieval — the loop these services are selling, written out.
- Research agents — the playbook for building rather than buying it.
- Trajectories — why the report is a summary written by the system you are judging.
- Hallucination & grounding — what a citation has to bind to in order to be worth anything.
- Unit economics — comparing meters that are not the same meter.
- Brave vs Exa vs Tavily vs Parallel — the search-API tier underneath all of this.