Drive-thru & restaurant ordering agents.
Voice AI running a drive-thru lane alone gets roughly 83% of orders right against about 87% for the standard lane — and roughly 95% when staff step in on the hard ones, which they do on about one order in five. Read those three numbers together and the product stops being a speech recogniser and becomes a routing decision: the thing you are building is the handoff, and the metric that moves the P&L is containment at an acceptable intervention rate, not accuracy.
Pick the metric the P&L actually moves on.
Accuracy is the number that gets quoted and the number that decides least. A 2026 study put AI-only drive-thru order accuracy around 83% against roughly 87% for standard orders, with staff-supported orders at about 95%; the 2025 Intouch Insight drive-thru study found a similar six-point gap and a hybrid figure of 97%. About 21% of AI orders still need employee support. Nothing in that spread tells you whether the lane makes money.
Four numbers do, and they should be on the same dashboard from week one:
- Containment. Share of orders completed without a crew member touching them. Yum has reported order containment above 90% across its voice-AI drive-thru deployments — that is the figure a rollout lives or dies on, because it is what converts into labour hours.
- Intervention rate and intervention cost. One in five orders needing help is fine if help costs eight seconds and catastrophic if it costs a crew member standing at a headset. The cost of an intervention is a design output, not a given.
- Order time. Wendy's reported at its 2025 Investor Day that FreshAI locations averaged 22 seconds less per order than human-staffed lanes. Seconds are cars per hour, and cars per hour is the whole business.
- Attachment. The same report cited a 15% lift in upsell attachment. An agent that always offers the combo is a revenue mechanism that does not get tired at 11pm — and one that offers it three times is a customer-satisfaction problem you will not see in accuracy.
Scale reference, so you can calibrate what "deployed" means: by late 2025 Wendy's had FreshAI in over 500 company and franchise locations; Taco Bell reported 890 lanes as of July, more than a tenth of its US footprint; White Castle runs SoundHound's "Julia" in around 100 drive-thrus, roughly 30% of the chain. These are production systems at chain scale, not pilots.
The escalation is the product.
The jump from 83% to 95% when staff step in is the most useful fact in the domain, and the usual reading of it — "the AI needs to get better" — is the wrong one. It says the hybrid system already works, and that its quality is set by how good the handoff is. A recogniser that improves from 83% to 86% moves the blended number barely at all. A handoff that goes from four seconds to one moves it a lot.
So design the escalation first, with an explicit trigger set and a defined crew experience:
# order/route.py — escalate on structure, not on ASR confidence def route(order, turn): if turn.unresolved_modifier: # "no ice, actually light ice" return ESCALATE if order.item_not_in_menu_version: # LTO ended Tuesday return ESCALATE if turn.repeat_count >= 2: # asked twice, still unsure return ESCALATE if order.total > LANE_MAX: # catering-sized, wrong channel return ESCALATE return CONTINUE
- Escalate on structure, not on a confidence score. "Model was unsure" is a weak signal that fires on accents and engine noise. "The caller has now amended the same modifier twice" is a strong one that fires on genuinely hard orders.
- The crew member must arrive with the order on screen. Handing over audio alone forces them to restart the conversation, which is exactly the eight-second cost you were trying to avoid. Ship the partial order, the disputed line, and nothing else. See escalation and warm transfer.
- Escalation must be available to the customer, not only to the model. "Let me get someone" spoken by a frustrated driver is a trigger. Never make it a menu option.
- Failing back to a human lane is the degradation path. When the agent is down, the lane runs as it did in 2019. Build that switch deliberately and test it, rather than discovering it during a lunch rush — the pattern is ordinary graceful degradation, and a lane-level kill switch belongs on the manager's screen, not in your deploy pipeline.
The hard part is the modifier grammar, not the transcription.
Teams arrive expecting the difficulty to be speech recognition and find it is combinatorics. A drive-thru order is a structured transaction against a catalogue where a single item carries a tree of options — size, preparation, substitution, removal, sauce, quantity-per-modifier — and customers express those in a grammar the menu does not use. "Number three, no onion, make it a large, and can you put the sauce on the side" is four mutations to one line item, delivered out of order, one of which ("make it a large") refers to a different field than it appears to.
- Ground every item in the live menu version. Limited-time offers appear and vanish weekly, availability differs by store and by daypart, and a model happily accepting an item that ended on Tuesday produces an order the POS will reject at the window. Retrieval against the current menu is not an optimisation here; it is the correctness boundary — the same discipline as grounding a product before checkout.
- Write through the POS, never around it. The point-of-sale system is the system of record for price, tax, promotion and the kitchen ticket. The agent's output is a POS transaction, and anything the POS cannot express is not an order — it is an escalation.
- Model the order as state, not as transcript. Each turn mutates a structured basket; confirmation is a read of that structure, not a replay of what was heard. This is what explicit dialogue state buys you, and it is why "repeat the order back" is cheap to get right and "remember what they said" is not.
- Let the confirmation board do the confirming. The lane already has a screen showing the order as it builds. It is a far better confirmation channel than speech, it costs no seconds, and it turns a verbal correction into a pointing gesture. Design for it rather than reciting the order aloud.
This is the worst audio you will ever ship against.
Everything that makes voice hard is present at once: a far-field microphone in an outdoor enclosure, engine and road noise, a car stereo, wind, weather, passengers talking over the driver, and a second lane running its own conversation three metres away. The speaker is also playing the agent's own voice into the same enclosure, which means acoustic echo cancellation is a correctness requirement rather than an audio-quality one — see full-duplex speech for why that gets harder, not easier, as models start talking while they listen.
- Test on the lane, not in the office. Word error rate on clean audio predicts nothing here. Build your evaluation set from real lane recordings across weather, dayparts and vehicle types, and keep it segmented — a model that is excellent at noon and poor at 2am is a staffing problem you can plan for only if you can see it.
- Passenger crosstalk is the signature failure. A back-seat voice ordering for themselves is not an interruption to suppress and not a correction to apply. Handling multi-speaker orders well is a distinguishing capability; handling them badly produces the "it added a drink I never asked for" complaint that dominates social media coverage of this category.
- Language mix is the norm, not an edge case. Many markets need code-switching mid-order rather than a language toggle at the start. Treat it as baseline capability and measure it separately; see multilingual voice agents.
Every widely-shared drive-thru AI failure video is one of three things: a phantom item from crosstalk, a modifier applied to the wrong line, or a loop where the agent asks the same question a fourth time. None of them is a transcription failure in the usual sense, and none will be caught by a word-error-rate gate. Build the failure taxonomy around those three.
The latency budget is set by the car, not the conversation.
In a call centre, latency is a politeness problem. In a lane it is a throughput problem with a physical queue attached: slow ordering backs cars up past the menu board, past the entrance, and eventually into the road, at which point drivers leave. Drive-offs are the metric nobody instruments until they have a bad week.
- Measure from vehicle detection, not from first speech. The clock the operator cares about starts when the loop detector fires. Greeting latency — detector to first word — is a distinct number from response latency and is usually the one customers describe as "it was slow".
- Budget the whole lane, not the model. The latency budget here includes the menu lookup, the POS write and the confirmation-board refresh. A 300 ms model in front of a 1.2 s POS round-trip is a 1.5 s agent.
- Barge-in is not optional. Drivers correct mid-sentence constantly, and an agent that finishes its sentence first loses more seconds than any model optimisation will recover. This is the single highest-value interaction behaviour in the domain — see turn-taking and barge-in.
- Do not spend seconds on upsell in a queue. Attachment lift is real, and it is worth less than throughput when four cars are waiting. Make the offer conditional on queue depth, which you already know.
Recording, disclosure, and the crew whose job you changed.
Two categories of obligation sit on this deployment and both are cheaper to handle before launch than after.
- Recording and voice data. Lane audio is recorded and retained for model improvement, which puts you in consent and biometric-privacy territory that varies sharply by jurisdiction — and a voiceprint used to recognise a returning customer is a materially different legal object from a transcript used to fix a menu bug. Decide what you retain, for how long, and whether any of it identifies a person; see recording, consent and redaction.
- Disclosure. A driver should know they are talking to a system, in the first utterance, in the channel they are using. The exemption everyone reaches for — "it is obvious" — gets weaker every quarter as the voices improve.
The operational half matters as much. The crew member's job changed from taking orders to supervising a queue of exceptions, which is a harder job requiring faster context switches, and it is performed during a rush. If the intervention experience is bad, the crew will route around the agent — taking over pre-emptively, or muting it — and your containment number will quietly decay for reasons no dashboard shows. Watch the manual-takeover rate per store, and treat a store whose takeover rate is climbing as a product bug rather than a training problem.
If you do one thing: instrument containment, intervention rate, greeting latency and drive-offs per store per daypart before you tune a single model parameter, and make the crew handoff carry the structured order rather than the audio. The 83% figure is not the ceiling you have to raise — it is the input to a routing decision, and the 95% hybrid number is proof that the routing is where the value is. Related: customer-support agents for the deflection-versus-confidence trade in its general form, voice failure modes for the catalogue of ways lanes break, and unit economics for turning seconds saved into a number a franchisee will act on.