Plan Review & Editing

9 min read

H24
Playbook · Agent UX & Human Interaction

Plan review and editing.

A plan screen looks like the humane alternative to twelve confirmation dialogs, and it is — but the trade it makes is rarely stated: you have replaced twelve informed decisions with one uninformed one. A plan is written before any tool has run, so every argument in it is a guess, and you are asking a user to authorise steps whose inputs do not exist yet. The design that works keeps those two jobs apart. Use the plan to let the user steer, and leave authorisation attached to the irreversible step at the moment its arguments become real.

STEP 1

Price the trade before you build the screen.

The motivation for plan review is sound and well documented: per-action gates produce confirmation fatigue, and the fatigue degrades the gate until it is a formality — the mechanism approval and confirmation UX works through in detail. Lifting the decision to the plan level is the obvious fix. It is also a trade with a term nobody writes on the whiteboard.

Do the arithmetic on both sides. A twelve-step task with a gate per step costs the user twelve interruptions, each carrying a fully-determined payload: the actual file, the actual recipient, the actual amount. The same task behind one plan approval costs one interruption carrying twelve intentions. Interruptions fall by eleven. Information per decision falls from "the exact bytes that will be written" to "a sentence about what will probably be written".

Here is the question that settles whether your plan is reviewable at all: name one step in it whose arguments are fully determined before the preceding step runs. For a plan that reads a database and then updates rows, or searches a repo and then edits files, the honest answer is usually none past step one. Everything downstream is parameterised by a result that does not exist yet. That is not a flaw in your planner — it is what plan-and-execute is for. But it means the plan cannot carry the approval, because the thing being approved is not in it.

The tell that a team has made this trade without noticing: the plan screen has an "Approve" button and no "Approve the first three steps" button. If the plan is genuinely an authorisation surface, partial authorisation is the common case and should be the easy one. If it is a steering surface, the button should say something closer to "Start" — and the gates should still be downstream, where the payloads are.

STEP 2

Editable has to mean more than a text box, and the edit has to survive replanning.

Read-only plans teach users to skim. An editable plan earns attention because the reader has a reason to look — and the reason has to be that changes stick. Four edit affordances cover almost everything users actually want, in descending order of how often they are wanted:

  • Delete a step. The most common edit by a wide margin, and the one users most often want for reasons the agent cannot infer — the step is somebody else's job, or it was done yesterday, or it touches a system that is mid-migration.
  • Constrain a step. Not "don't search", but "search only packages/api". This is where the user's unexpressed context enters, and it is worth more than any amount of prompt engineering because it is specific to this run.
  • Pin a value. The user knows the target branch, the cost centre, the recipient. Letting them fix it in the plan removes a later clarifying question and an entire class of wrong guess — see clarifying questions for when to ask versus assume.
  • Reorder. Rarely needed and expensive to honour, because order usually encodes a data dependency the user cannot see. Offer it last, and reject it with an explanation rather than silently re-sorting.

Now the part that decides whether any of this is real: the agent will replan. Something will come back differently and the plan will be regenerated mid-run. If the step the user deleted reappears, or their constraint is dropped because it lived in the plan object rather than in the instruction, you have taught them that the control is decorative — and that lesson generalises to every control in your product. Edits have to be promoted out of the plan and into the run's standing constraints, where the planner reads them again on every regeneration. A deletion is a prohibition, not a diff.

STEP 3

Separate the two promises a plan makes, because users only read one.

A plan silently asserts two different things, and the UX almost always shows the first while the user is pricing the second.

  • Scope — the set of systems, files, records and people this run will touch. This is the reachable set, and it is the thing task scope argues you should state explicitly rather than let the reader infer from a list of verbs.
  • Autonomy — which of these steps will come back and ask, and which will not. Users overwhelmingly assume they will be asked again about anything serious. Most plan screens make no such promise, and most agents behind them do not keep one.

Make both legible per step. A three-value marker is enough — auto, ask, ask-with-payload — and the third is the one that matters, because it is the promise that the actual arguments will be shown before the irreversible thing happens. Then surface the aggregate, because it is the number users actually want and nobody displays it: this plan has nine steps and will interrupt you twice. A user who knows they will be asked twice reads the plan as steering and gets on with their day. A user who does not know reads it as a contract and reads it wrong.

Which steps get which marker is a consequence question, not a plausibility one, and the tiering is already worked out: gate on reversibility and blast radius, not on how dramatic the verb sounds. Undo and reversibility is the input — a step you can cleanly undo does not need an ask, and a step you cannot should never be auto no matter how routine it looks.

STEP 4

Plans drift. Show the diff, and re-ask only when the diff crosses something.

Step three returns a surprise, the planner regenerates, and the run is now executing a plan the user never saw. This is the failure mode that makes plan approval worse than no approval, because the user believes they reviewed something and the artefact they reviewed no longer exists. Every approval is a claim about a world that has since moved — the general treatment is time-of-check to time-of-use, and a plan screen is the longest such window most products ship.

The fix is a plan diff with a re-ask rule, and the rule has to be mechanical or it will not survive contact with a deadline.

  • Keep the approved plan as an artefact, not a rendering. You cannot diff against a string you regenerated. Store the plan as structured steps with scope annotations so "did this change" is a computation rather than a judgement.
  • Re-ask when the diff adds a system, widens the scope, or adds an irreversible step. Three conditions, checkable in code. Everything else — reordering, retries, a step that turned out unnecessary — gets shown in the timeline and does not stop the run.
  • Show the drift where the user already is. A banner on the run view that says "the plan changed: one new system" is read. An email is not, and a modal that arrives while they are in another tab is an interruption problem of its own.
  • Count the drift and show it afterwards. "This run executed eight of the nine steps you approved, plus two added at step four" is an honest post-run summary, and it is the sentence that calibrates how much the next plan is worth reading.
STEP 5

Four cases where a plan screen is the wrong product decision.

Plan review has become a default, which means it is now being added to agents that do not benefit from it. Each of these is a case where the screen costs more than it returns.

  • The task is shorter than the plan. If generating and reading the plan costs more wall-clock than executing it, you have built a speed bump. Below roughly three steps, skip it: run and show an undo.
  • The run is read-only. Nothing is irreversible, so there is nothing to authorise, and the plan's only remaining job is steering — which a good streaming view does better and without a gate. Save the screen for runs with effects.
  • The user cannot evaluate the plan. A plan referencing six internal services is unreviewable by someone who knows four of them, and asking anyway manufactures a rubber stamp while transferring accountability to the person least able to carry it. This is the same arithmetic that makes a 98%-approval review queue a broken control in review queues for agent output. The honest answer is a narrower agent, not a longer plan.
  • The task recurs unchanged. If the plan is the same every Tuesday, you are asking a human to re-approve a procedure. Promote it to a saved, versioned procedure that is reviewed once when it changes — which is what user-authored skills and spec-driven agent development are both reaching for from different ends.

Where it does earn its place, it earns it twice: a plan is the cheapest point in the run to correct a misunderstanding, and it is the natural place to raise the autonomy ceiling over time, because an edit rate that falls as trust accrues is exactly the evidence progressive autonomy asks for.

STEP 6

Instrument three numbers, and be willing to delete the screen.

Plan review is unusually easy to evaluate, because the user's behaviour tells you directly whether the screen is doing anything. Three rates, and each one has an action attached.

  • Edit rate. The fraction of plans a user changes before starting. This is the whole justification for the feature: if it sits near zero, users are approving without reading and you have bought one rubber stamp in exchange for the twelve gates you removed. Below a few percent, either the plans are genuinely right — in which case stop showing them and keep the gates on irreversible steps — or they are unreadable, which you can distinguish by time-on-screen.
  • Post-approval replan rate. How often the plan changed materially after it was approved. High rates mean the planner is guessing at plan time, so the review is theatre; the fix is to plan less far ahead, not to plan better. A planner that emits three confident steps and then re-plans is more honest than one that emits nine speculative ones.
  • Post-completion intervention rate. Undos, rollbacks and apologies after a plan was approved. This is the only number that tells you whether plan approval protected anybody, and it is the one teams do not connect back to the plan screen. If it does not fall when you ship plan review, the screen is not a control.

Segment all three by user, because the aggregate hides the pattern that matters: new users edit plans and experienced users do not, and the experienced users are the ones running the consequential tasks. A feature whose protective value decays exactly as the stakes rise needs the gates kept downstream, which is the argument this page has been making from the top.

Ship it in this order. First add the per-step ask / ask-with-payload markers and the interruption count, because they cost a day and they stop the plan from being read as a promise it is not making. Then make deletions and constraints into standing constraints that the planner re-reads on every regeneration — this is the change that turns the screen from decoration into a control, and it is a backend change, not a UI one. Then instrument edit rate. If after a month it sits near zero for your experienced users, delete the plan screen for them and keep the downstream gates: you will have removed an interruption and lost nothing, which is the outcome the arithmetic predicted.