Clarifying Questions

A30
Concepts · Agentic AI Explained

Clarifying questions.

The agent that asks whenever it is unsure is not the safe one — it is the one people stop delegating to, because every question spends a human's attention to buy certainty the agent could usually have bought itself. Confidence is the wrong trigger. Ask when the ambiguity would change the work and the wrong branch is expensive to undo; otherwise state the assumption you are proceeding under, and make it cheap to reverse.

STEP 1

A question is an action, and it has a price.

It is tempting to treat asking as the conservative default: worst case, you annoyed someone. That accounting is wrong in three places, and the third is the one that kills products.

  • Latency, which is not symmetric. An interactive agent pays seconds for a question. A background or scheduled run pays whatever the gap is until a human next looks — often hours, sometimes the rest of the run, because the question arrives in a queue nobody is watching. The same question costs three orders of magnitude more depending on where the loop is running.
  • Attention, which is the scarce input. Answering costs the human a context switch back into a task they delegated precisely to stop holding. Two questions during a ten-minute task can exceed the cost of doing the task by hand — which is the comparison the user is actually making.
  • Trust, which is asymmetric and slow to rebuild. Users read frequent questions as "this thing cannot be left alone", and they respond by shrinking what they delegate. An agent that asks about everything converges on being used for nothing, and the telemetry shows it as falling usage rather than as a clarification problem.

The inverse failure is real too, and it is the one goal drift describes: an under-specified request gets silently resolved into an easier adjacent task, competently executed, and returned as a good answer to a question nobody asked. The point is not that asking is bad. It is that "am I uncertain?" does not distinguish the two cases, and something else has to.

STEP 2

The gate is two questions, and neither is about confidence.

Model uncertainty is a poor trigger because it is both badly calibrated and beside the point — see uncertainty & calibration. A model can be genuinely unsure about something that does not matter, and serenely confident about the one reading that will cost you a production database. Gate on consequence instead, with two tests applied in order:

  • Is the ambiguity load-bearing? Do the plausible readings lead to materially different work? "Should the report be Markdown or HTML" usually does not — pick one, say which. "Delete the stale records" when stale could mean thirty days or three years absolutely does.
  • Is the wrong branch expensive to reverse? A wrong guess you can undo with an edit is a draft. A wrong guess that sends an email, moves money, or drops a table is not a draft. This is the same axis human in the loop uses for approval gates and blast radius uses for scoping.

Both true, ask. Load-bearing but reversible: proceed on a stated assumption and surface it — this is the large middle cell, and defaulting it to "ask" is the single most common design error. Irreversible but not load-bearing: you do not have a question, you have a confirmation, which is a different interaction with a different cost. Neither: proceed silently.

Note what the gate does to a background agent. Because the top-left cell is the only one that stalls the run, an unattended run should be specified so that cell is empty before it starts — the irreversible, genuinely ambiguous decisions are resolved in the brief, not discovered at 03:00.

STEP 3

When you do ask, ask once, early, and answerably.

Most of the pain attributed to clarifying questions is really bad question design. Four rules cover it.

  • Never ask what you can look up. A question the agent could have answered with a tool call is a bug, not a courtesy — "which database?" when there is exactly one connection configured is the canonical case. Resolve from the environment first; ask only about intent, which is the one thing not written down anywhere.
  • Batch at the front. Three questions asked together cost one interruption; three asked across twenty minutes cost three, and the third arrives after the human has moved on. Plan far enough ahead to know what you need before you start — which is a large part of what planning is for, per planning & termination.
  • Carry a default. "I'll treat stale as older than 90 days unless you say otherwise" is answerable by silence and by one word. An open "what did you mean?" makes the human do the work of enumerating the options, which is the work you were delegated.
  • Ask for the decision, not for the specification. Offer two or three concrete branches with their consequences. Users are far better at rejecting a wrong option than at producing a complete one, and the exchange terminates in a single turn.

Protocol support exists for exactly this shape: MCP's elicitation lets a server request a specific, schema-typed piece of input mid-call rather than failing or guessing — see sampling & elicitation. The mechanism is not the hard part; the policy about when to invoke it is.

STEP 4

The alternative is a stated assumption, and it needs to be visible.

For the reversible-but-load-bearing majority, the move is to decide and declare: proceed under a named assumption, put it where the result is read, and make reversing it cheap. An assumption stated in the first line of the output costs the reader two seconds and costs the agent nothing; the same assumption made silently is indistinguishable from a hallucination when it turns out wrong.

  • Record assumptions as structured output, not as prose buried mid-answer — they belong next to the result, and in the trace, so a reviewer can scan them without re-reading the work. This is the same argument agent UX patterns makes about showing intermediate state.
  • Make the reversal one step. "Re-run with --since=3y" is a stated assumption with an exit; "regenerate the whole report" is not.
  • Measure the right thing. Clarification rate is a vanity metric that can be driven to zero or one by prompt tone alone. The numbers that matter are rework caused by an unstated assumption and runs abandoned at a question — the first tells you that you are asking too rarely, the second that you are asking too often, and a healthy system has a little of both.

Write the policy into the system prompt as a rule about consequences, not about confidence: resolve what the environment can answer, ask once up front only when a genuinely ambiguous decision is also hard to undo, and otherwise proceed on an assumption stated in the output. Then go read your last fifty questions and delete the ones that a tool call would have answered — in most deployments that is the majority of them, and removing them is the cheapest trust you will ever buy.

Related: human in the loop for approval gates, uncertainty & calibration for why confidence misleads, and waiting & latency UX for what the human sees while the agent decides not to ask.