First Run & Onboarding

10 min read

H11
Playbook · Agent UX & Human Interaction

First run: the session that sets a prior your product spends six months arguing with.

A user's first session with an agent produces a durable estimate of what it can do, and every later interaction is read through that estimate rather than replacing it — so a first run that showcases the ceiling manufactures a disappointment on task two, and a first run that fails once costs you more trust than ten failures a month later would. The onboarding that actually works demonstrates the boundary instead of the capability: show what the agent will refuse, show it being steered mid-task, and pick a first job you can guarantee over one that impresses.

STEP 1

The first session sets a prior, and priors are updated slowly and asymmetrically.

People do not build a fresh model of an agent's competence on every use. They form one early, from very little evidence, and then spend a long time interpreting new evidence in its light. This is ordinary human behaviour and it is unusually consequential for agents, because the thing being estimated — "what can I hand this?" — is invisible, has no natural edges, and is the single decision that determines whether the product gets used at all.

  • An early failure is expensive because there is no track record to absorb it. The twentieth session's mistake lands against nineteen successes and reads as an exception. The first session's mistake is the data set. Users who see one confident wrong answer in session one routinely conclude the agent "makes things up" and never revise it, even as the model behind it improves.
  • An early over-performance is also expensive, in the opposite direction. A demo-grade first task teaches a boundary that is too wide. The user's second request sits just outside it, fails, and now they have learned two things at once: that the agent is unreliable, and that they cannot tell in advance which side of the line they are on. The second lesson is the one that kills usage.
  • The estimate people form is about the boundary, not the average. Nobody adopts a tool on its mean quality; they adopt it once they can predict which tasks are safe to delegate. Predictability beats capability at this stage, and a narrower agent that is legible is adopted faster than a broader one that is not — which is the same argument designing for trust makes about calibration.
  • Recovery is not free even when it works. Winning back a user who wrote the agent off in week one usually requires a new reason to try again — a colleague's recommendation, a visible release — and you get very few of those.

The design consequence is uncomfortable but clarifying: the first run is not a marketing surface. Everything about it should be optimised for the accuracy of the mental model it produces, not for how good the product looks while producing it. Those two goals conflict more often than they align.

STEP 2

The capability tour teaches the ceiling. Teach the edge instead.

The default onboarding — a carousel of impressive things the agent can do, or a set of suggested prompts hand-picked to succeed — is optimised to communicate range. It communicates it accurately, and that is the problem: range is the top of the distribution, and the user calibrates to the top.

  • Show a refusal on purpose, early. An agent that declines something clearly out of scope, and says why in one sentence, teaches more about where the line is than five successes do. It is also the single cheapest credibility signal you have, because a tool that says no is a tool whose yes means something.
  • Show an uncertain answer as well as a confident one. If the agent hedges when the evidence is thin, the user learns that the hedge carries information — and starts trusting the confident answers more. If everything arrives at the same volume, they learn nothing and default to verifying all of it, which is the failure mode where the agent creates work.
  • State the boundary in one line, in the agent's own voice, at the top. "I can read your docs and draft replies; I can't send anything or touch billing." Specific, checkable, and worth more than a features page. See transparency and explainability for how much of this belongs in the interface rather than the documentation.
  • Do not let suggested prompts be a highlight reel. If the starter set is chosen for impressiveness, it is a lie with a good conversion rate. Choose the three tasks that are most reliable and most common; if those two lists do not overlap, that is your roadmap talking.
  • Never promise autonomy you have not turned on. Onboarding copy that describes what the agent will eventually do reads, to a first-time user, as what it does now. That gap is where the first disappointment comes from, and progressive autonomy is the mechanism for widening the boundary honestly over time.
STEP 3

Pick a first task you can guarantee, and make it the user's real work.

The first-run task is the highest-leverage product decision in the whole onboarding, and it is usually made by whoever wrote the empty state. Two properties matter, and they trade against each other less than you would expect.

  • Reliability first, at a level you would bet the account on. Whatever the first task is, it should be the path with the best measured success rate you own — not the newest feature, not the one the demo used. If your evals cannot tell you which path that is, you are not ready to design a first run.
  • It has to be their work, not a sandbox. A tutorial task with fake data teaches the mechanics and nothing about applicability, and users discount success on toy input almost completely. Getting a real, small, verifiable result on their own material is the whole event.
  • Small and verifiable beats large and impressive. The user needs to be able to check the output in under a minute. A result they cannot verify does not create trust — it creates an obligation to review, which feels like work and reads as risk.
  • Reduce the setup between arrival and first result to as close to nothing as you can. Every connector, permission and configuration step before the first output is a place to leave, and each one is being paid for by a user who has not yet seen anything work. Ask for the minimum the first task requires and defer the rest to step 4.
  • If the agent will be slow, say what it is doing. A first run that sits silent for ninety seconds is interpreted as broken far more often than as thorough. Streamed intermediate steps are worth more here than anywhere else in the product, because the user has no basis yet for patience — see streaming and partial output.
STEP 4

Permissions granted at first run are not consent, and bundling them is a dark pattern by accident.

The standard flow asks for every scope the agent might need before it has done anything. This is convenient for engineering and indefensible as consent: the user is being asked to authorise capabilities whose consequences they cannot yet estimate, at the moment they have the least information they will ever have, while motivated to get past the screen.

  • Ask at the moment of use, tied to the task that needs it. "To draft this reply I need read access to the thread" is a request a user can evaluate. A checklist of nine scopes on screen two is not, and the approval it produces is worth nothing when something goes wrong.
  • Separate read from write, always, and start read-only. The most valuable first-run configuration for most agents is one that cannot change anything. It removes the entire class of first-session disasters and it costs a capability the user was not going to use yet anyway.
  • Show the consequence, not the scope name. "Can send email as you" is the honest rendering of a permission string that reads as harmless. This is the same discipline as approval and confirmation UX, applied before there is any history to reason from.
  • Make the first grant visibly revocable, and put the revocation where they are. Knowing you can take it back is what makes granting it reasonable. A permissions page three menus deep does not provide that assurance to someone who has used the product for four minutes.
  • Log first-run grants as their own class and review them. If most users are approving the maximum scope in under five seconds, your flow is not obtaining consent, it is harvesting clicks — and that is a governance finding, not a conversion win.
STEP 5

Teach the repair, because the user's first correction is the moment they decide whether this is collaborative.

Every agent goes off course. Whether that is fatal depends almost entirely on whether the user knows what to do about it — and the first run is the only time you have their full attention to show them.

  • Demonstrate steering mid-task at least once. Let the user interrupt, add a constraint, and watch the agent incorporate it without starting over. A user who has done this once has a fundamentally different relationship with the tool than one who has only watched it run to completion.
  • Make undo visible before it is needed. The presence of a working undo changes what people are willing to let the agent attempt, which means it raises the value of the first session rather than merely insuring it — the argument laid out in undo and reversibility.
  • Show the trace for the first result, whether or not you show it later. What was read, what was called, what was decided. The user is building a model of how the thing works and this is the only artefact that supports it. After a few sessions most people stop looking, which is fine — it did its job.
  • Handle the first failure as a designed path, not an error state. If the first task fails, say what failed, what it did not touch, and offer the narrower version that will work. A first session that fails gracefully and recovers can produce better calibration than one that succeeds silently. Designing for failure is at its highest leverage here.
  • Make correction cheap and obviously wanted. The first time a user edits the output, they are teaching you and testing whether the product wants to be taught. If the edit vanishes into a text box with no acknowledgement, they learn the answer.
STEP 6

Measure the first session separately, because your aggregates are built to hide it.

Blended quality metrics are dominated by your engaged users, who have already learned the boundary, formed the workarounds and stopped asking for the things that fail. The population you most need to see — people in their first ten minutes — is a rounding error in every number on the dashboard.

  • Report first-session success as its own metric, defined as a verified useful result on real data. Not "completed onboarding", not "sent a message". If you cannot define the result, you have not defined the first-run task.
  • Watch time-to-first-useful-result, at the tail. The p90 is the experience of the users most likely to leave, and it is the number the median comfortably conceals.
  • Track the return rate at day two and day seven for cohorts split by what happened in session one. The gap between "first task succeeded" and "first task failed" cohorts is the honest price of a bad first run, in a currency your leadership already reads.
  • Read the first ten transcripts of every week by hand. Not sampled, not summarised. The specific place where a new user's phrasing missed the agent's actual capability is visible in thirty seconds of reading and invisible in any aggregate, and this is the cheapest research in the building.
  • Re-run onboarding as a fresh user after every capability change. The boundary you describe on screen drifts out of date silently, and the person least likely to notice is the one who built it. Put it in the release checklist next to rollout and versioning.

If you do only one thing this week: open your onboarding as a new user, write down the boundary a reasonable person would infer from it, then compare that against your eval results for the same tasks. The distance between those two lines is the disappointment you are shipping on purpose, and closing it usually means removing a promise rather than building a feature. Onboard to the edge of the agent, not to its ceiling.

Related: designing for trust for the calibration argument in full, progressive disclosure for how much to show and when, and agent UX patterns for the vocabulary.