AI Blog

OpenAI vs Gemini vs Perplexity vs Exa: the research API sells you the loop

A search API returns documents and leaves the agent loop in your process. A research API takes the loop, and that is the trade — you stop paying to orchestrate and you stop being able to instrument. The axis nobody tables is what a citation is: three of these four hand back a bibliography the model assembled, and one binds grounding to a field in a schema you defined, with a confidence. Pick on that, not on report quality.

By Agentic AI Wiki 14 min read

If your research output is going into a record rather than in front of a person, only one of these four can tell you which source supports which field — and that single difference outranks every comparison of report quality you will read. A search API returns documents and leaves the agent loop in your process. A research API takes the loop, which is the product and the price at the same time: you stop paying to orchestrate, and you stop being able to instrument.

At a glance

Four services that accept a question and return findings, with the loop on their side of the wire.

ServiceInterfaceDefault corpusBilling unit
OpenAI deep research (o3-deep-research, o4-mini-deep-research)Responses API, background modeNone by default — you must attach web_search_preview and/or an MCP serverTokens: $10/M in, $40/M out
Gemini Deep Research agentInteractions API, async only, paid tier onlyGoogle Search, URL Context, code executionPer task: roughly $2 standard, $5 Max
Perplexity Sonar Deep ResearchChat-completions shapedPerplexity's own indexFive meters, listed below
Exa Research / Exa AgentREST, sync or asyncExa's neural indexPer request: roughly $12–15 per 1,000
Research APIs across four axes A four-by-four matrix. OpenAI: bibliography-level citations, private corpus via MCP, token billing, some loop control through tool choice. Gemini: per-claim sourcing, private corpus via MCP and file search, fixed per-task billing, no loop control. Perplexity: bibliography-level citations, web-only corpus, five billing meters, no loop control. Exa: field-level grounding with confidence, web-only corpus, fixed per-request billing, effort modes. What separates the four Citation granularity Private corpus Bill predictability Loop control OpenAI Bibliography Via MCP Token-metered Tool choice Gemini Per claim Via MCP + files Per task Fixed Perplexity Bibliography Web only Five meters Fixed Exa Field-level Web only Per request Effort modes Strong Partial Weak or absent
Three of the four columns separate these services cleanly. The fourth is where they are all the same.

Prices are list prices at the time of writing and all four have moved in the last six months; treat the ratios as durable and the absolute figures as perishable.

What you are actually buying

Where the research loop lives Two panels. On the left, over a search API, your process owns the plan, search, read, judge and synthesise steps and calls the search API once per query, producing five trace spans you own. On the right, over a research API, your process sends one question and receives one report plus a URL list, while plan, search, read, judge and stop all run inside the vendor boundary. Search API — you own the loop your process Plan the sub-questions Issue a query Fetch and read Enough evidence? Synthesise search API Five spans in your trace. Every stopping decision is yours. Research API — the vendor owns it your process One question Vendor boundary plan · expand · fetch · read · judge · stop query budget, depth and stopping rule set by the service, not by your request One report, plus a list of sources One span. The report is a summary written by the system you would debug.
The boundary moved. Everything else about these products follows from where it landed.

Build a research agent on a search API and you write the loop: decompose the question, issue queries, fetch and read, decide whether the evidence is sufficient, issue more queries, synthesise. It is perhaps three hundred lines, it is tedious, and it is where all your cost and all your failure modes live. Every one of those steps is also a span in your own trace, scored by your own evals.

A research API deletes those three hundred lines. In exchange, the plan, the query expansion, the read/skip decision and the stopping rule all move behind an API boundary you cannot see through. What comes back is a report and a list of sources. You can evaluate that report end to end — and end to end is the only way you can evaluate it, because the trajectory that produced it is not yours.

That is a legitimate trade and frequently the right one. It stops being right the moment someone asks why the answer is wrong, because the honest reply is that the report is a summary written by the system you are trying to debug. Everything below is about which of the four gives you the most to work with when that happens.

The axis nobody tables: what a citation is

All four advertise citations. They mean two different things by the word, and the gap between them is the single largest difference in this comparison.

A bibliography is not a claim-level guarantee

OpenAI's deep research models and Perplexity's Sonar Deep Research both return prose with a set of URLs attached. The URLs are real and were really visited. What they are not is a binding between a particular assertion and the evidence for it. If the report says a threshold is 40 milliseconds and lists nine sources, you cannot programmatically determine which of the nine — if any — contains that number. A human can check it by reading nine pages. A downstream system cannot check it at all.

Gemini sits in the middle. The Deep Research agent provides granular sourcing for claims rather than only a document list, which is materially better for a reader and still not a machine contract: you get an attribution, not a schema you asked for with a value in it.

Field-level grounding is a different product

Exa is the outlier and it is the reason this comparison exists. You pass an outputSchema; you get back output.content matching that schema and output.grounding carrying citations and a confidence per field. That is not a nicer citation format. It is the difference between a document you read and a record you can write into a database with a provenance column already populated — the thing that the preceding post on agent-supplied values in authoritative records argues you should never ship without.

The decision rule is short. If a human reads the output and acts on it, a bibliography is fine and you should optimise for report quality. If the output populates fields — a CRM record, a pricing table, a due-diligence sheet, an eval set — you need grounding bound to the field, and today that narrows the field to one. Retrofitting field-level grounding onto a prose report by asking a second model to align claims to sources is a known-bad pattern: you have added a second ungrounded generation step and called it verification.

Whose corpus, and whether you can extend it

The second real split is whether the loop can be pointed at anything of yours. It is a cleaner divide than the marketing suggests.

  • OpenAI requires you to choose one. A deep research request must attach at least one data source: web_search_preview, a remote MCP server, or both, and code_interpreter is available alongside. The MCP server is not an ordinary one — the deep research models expect a specific shape, a search tool returning top-k identifiers and a fetch tool resolving those identifiers to documents. That constraint is doing real work: it forces your corpus into the same retrieve-then-read interface the web tool uses, which is why the loop can treat both sources identically.
  • Gemini supports remote MCP servers and file search, but not custom function calling. You can restrict the agent to the web (Google Search plus URL Context), to private sources (file search plus your MCP servers), or to a mix. What you cannot do is hand it an arbitrary function — so a research task that needs to call your pricing service mid-loop is out of scope, and the workaround is to wrap the service as an MCP server with search/fetch semantics it does not naturally have.
  • Perplexity and Exa are the open web, full stop. Both are excellent at it and neither pretends otherwise. If your research question needs your wiki, your data room or your ticket history, these are the wrong tier of product and no amount of prompt context substitutes — you are asking the model to reason about documents it cannot go and get.

So the build-versus-buy question is not about quality. It is whether your research corpus is public. If it is, buy; the vendors are better at web research than your loop will be. If the answer lives half inside your company, you are choosing between OpenAI and Gemini on MCP ergonomics, or you are back to writing the loop over an agentic retrieval layer of your own.

The money: four incommensurable meters

How many meters each service runs Four columns of billing meters. OpenAI runs two: input tokens and output tokens. Gemini runs one headline per-task price with a search-grounding line item underneath it. Perplexity runs five: input tokens, output tokens, citation tokens, reasoning tokens and search queries. Exa runs one: a fixed price per request or per run mode. Meters you are billed on OpenAI input tokens output tokens 2 meters, uncapped Gemini per task search $14 / 1K 1 headline, 1 line item Perplexity input tokens output tokens citation tokens reasoning tokens search queries 5 meters, two invisible Exa per request / run 1 meter, capped by mode Accent marks a bill you can forecast before the run; the forecast costs you the ability to let one question run long.
The number of meters predicts how badly your first month's bill will surprise you.

Written out, the list prices at the time of writing are:

OpenAI   o3-deep-research     $10 / M input      $40 / M output
Gemini   deep-research        ~$2 / task (standard), ~$5 / task (Max)
         underneath: Gemini 3.1 Pro at $2 / M in, $12 / M out
         plus Google Search grounding at $14 / 1K queries
         (~80 queries standard, ~160 Max; implicit caching
          covers 50-70% of input tokens)
Perplexity sonar-deep-research  $2 / M input     $8  / M output
         + $2 / M citation tokens
         + $3 / M reasoning tokens
         + $5 / 1K search queries
Exa      research w/ schema   ~$12-15 / 1K requests
         agent fixed modes    $0.012 (Minimal) - $1.00 (X-high) / run

These cannot be compared without fixing loop depth, and loop depth is the vendor's decision, not yours. A token-metered service has no ceiling: a question that triggers forty searches costs four times a question that triggers ten, and nothing in your request said how many to run. That is the whole difference between OpenAI's meter and Exa's.

  • Gemini's per-task price is the most honest artefact in the group, precisely because the line items are published underneath it. Eighty grounding queries at $14 per thousand is $1.12 of search inside a $2 task, which tells you immediately that the inference is nearly free and the retrieval is the product. It also tells you the fixed price holds only while the query budget does.
  • Perplexity's five meters are the most unpredictable, and the two unusual ones are the reason: citation tokens and reasoning tokens are volumes you neither specify nor see in advance. The prices are individually low. The variance is not.
  • Exa's fixed modes buy predictability with a capped effort level. Minimal to X-high is a dial you set per run, which is the right shape for a production pipeline and the wrong shape for a question that deserves to run long because it turned out to be hard.
  • Nobody bills for the retry. When a run returns a report that fails your quality gate, you pay again. Fold your gate's pass rate into every figure above before comparing anything — a service at 70% of the price with a 60% pass rate is more expensive, and this is the same arithmetic as cost per successful task everywhere else.

They are all jobs, not calls

Every one of these runs for minutes, and two of them refuse to pretend otherwise: Gemini's Deep Research agent is asynchronous only, and OpenAI's deep research models are meant to be driven in background mode because a synchronous request will not survive the run. Exa offers both shapes and the async one is the one you want above the cheapest modes.

The operational consequence is the part teams discover late. You need somewhere to persist the job handle, a poller or a webhook, a retry that does not re-run a completed job, a place to put a partial result when the run fails at minute nine, and a user-facing story for the wait. That is the durable-state infrastructure you were hoping to skip by buying the loop — buying the loop removes the orchestration of research steps, not the orchestration of the job.

Two smaller things that bite:

  • The corpus is untrusted input and the loop reads it autonomously. A research agent's entire job is to fetch and read pages written by strangers, and three of these four will also follow links. Whatever comes back is data, not instruction, and the fact that the vendor owns the loop does not make it their problem — the report lands in your context window. See prompt injection.
  • Reports are long, and length is a downstream cost. A 6,000-word report is 8,000 tokens in every subsequent turn of whatever agent consumes it. Ask for the schema you need rather than the essay, where the service lets you.

When to pick which

SituationPickWhy
Output populates structured fields in a recordExaThe only field-level grounding with a confidence; everything else needs a second ungrounded pass to fake it
The answer lives partly in internal systemsOpenAI or GeminiBoth take remote MCP servers; OpenAI forces a search/fetch shape, Gemini forbids custom functions
A human reads the report and decidesGemini or OpenAILongest, best-synthesised output; per-claim sourcing on Gemini is the better reading experience
High volume, fixed budget per itemExa or GeminiPer-request and per-task pricing are the only two that let you forecast; token meters cannot be capped
Freshness matters more than depthPerplexityPurpose-built research API over a live index, and the fastest of the four to first result
You need the trajectory, not just the answerNone of themWrite the loop over a search API — this is the case the category cannot serve

FAQ

Is a research API just a search API with a prompt on top?

No, and the difference is the loop rather than the prompt. A search API answers one query; a research API plans a sequence of queries, decides what to read, decides when it has enough, and writes the synthesis. You are buying the stopping rule as much as the search.

Which one produces the best report?

For long, hard, ambiguous questions read by a person, OpenAI and Gemini are consistently the deepest, and Perplexity is the fastest to something adequate. That ranking is also the least durable thing in this post — it has changed twice this year, and it is the wrong axis to architect on.

Can I get field-level citations out of the other three?

Not natively. The common workaround is a second model pass that maps report sentences to the source list, which produces confident alignments that are themselves ungrounded. If you need the guarantee, get it from the service that computes it during retrieval, not from a post-hoc aligner.

Do any of them read my private data?

OpenAI and Gemini can, through remote MCP servers — OpenAI requires a search/fetch interface specifically, and Gemini additionally offers file search while not supporting custom function calling. Perplexity and Exa search their own web indexes only.

How should I evaluate one of these?

Build a set of questions whose answers you already know and that are not answerable from a single page, then score the answer, the recall of sources you know exist, and the rate at which the report asserts something no cited source contains. Only the third is unusual and it is the one that separates the four. See why agent eval is hard.

Is it cheaper to write my own loop?

On a per-task basis, usually yes, once you already have a search API contract and a harness. On a total basis, usually no, until volume is high enough that your engineering time amortises — and the crossover is not about cost per call, it is about whether you need the trajectory for eval or debugging. That need, not the price, is the reason to keep the loop.

Further reading

On this wiki:

Project sources: