Taking a card payment over voice: the agent has to stop being the interface.
Every descoping technique the contact-centre industry spent twenty years building works by removing a listener from the audio path — and a voice agent is not a listener on the call, it is the call, which is why the standard playbook does not port. Get this wrong and one spoken card number lands simultaneously in a transcript, a model prompt sent to an inference provider, tool-call arguments, a trace span, an eval fixture and a CRM summary: six copies, six retention regimes, and at least two of them holding data that PCI DSS says may never be stored after authorisation. The design that works accepts a thing teams resist — for the length of the payment step, the agent hands the microphone to something deterministic and waits for a token.
Why the standard descoping pattern does not transfer.
The mature solution for human agents is DTMF suppression: the caller types their card on the keypad, a component in the telephony path intercepts the tones, and what reaches the agent's headset, screen and the recorder is flat tones or asterisks. The human stays on the line the whole time — reassuring, handling questions, watching a masked field fill up — and cardholder data never enters the agent environment. Vendors quote this as the move that drops a contact centre from SAQ D toward SAQ A, and the PCI Security Standards Council's own telephone-payments guidance treats suppression as a legitimate scope-reduction technique, with the caveat that eligibility depends on the specific implementation rather than on the product label.
That pattern has a hidden precondition: the listener you are excluding is not the thing driving the conversation. Remove a human's ability to hear sixteen digits and the call continues normally. Remove a model's, and you have removed its input. A cascade voice agent's entire perception of the call is the transcript; a speech-to-speech model's is the audio itself. There is no seat at the table for the agent during digit capture that does not put cardholder data into the agent's context.
Then count the copies, because this is the part that surprises people who have run a compliant contact centre for years. One spoken card number in an agentic stack produces:
# artefacts created by one spoken PAN, and who holds them 1. call recording your storage, your retention rule 2. ASR transcript often the speech vendor's too 3. model context/prompt the inference provider's request logs 4. tool-call arguments your orchestrator, your app logs 5. trace span your observability vendor, 30–90d 6. CRM summary + eval set indefinite, and copied onward # PCI DSS: PAN must be unreadable wherever stored; # CVV/CVC and PIN must not be stored AT ALL post-auth. # Rows 2-6 did not exist in the human-agent design.
Row 3 deserves its own sentence. Sending a card number to a third-party model provider is a disclosure of cardholder data to a service that is almost certainly not in your PCI scope, was not assessed, and may retain request payloads for abuse monitoring. No masking applied after the response helps; the data left when the request left. The five artefacts below the recording are the ones recording, consent and redaction names in general terms — this page is the specific case where they are not merely inconvenient but disqualifying.
One framing that saves a lot of argument in design review: for the payment step, treat the model as an untrusted third party that happens to be inside your stack. You would not pipe raw PAN and CVV to an unassessed external vendor to decide what to say next. That is what a naive voice-payment flow does, with the vendor's name in your inference bill.
Three architectures, and the trade you are actually making.
There are exactly three shapes that keep cardholder data out of the model, and they differ in what they cost the caller rather than in how compliant they are. Pick on caller experience and completion rate; all three descope.
- Suppression at the media layer, agent stays on the call. The SBC or media server intercepts DTMF, forwards flat tones to every downstream consumer — recorder, ASR feed, model audio input — and sends the digits to a payment service over a separate channel. The agent receives only field-level events:
pan_captured(last4, brand, luhn_ok),expiry_captured,cvv_captured. This is the only shape where the conversation stays continuous, the agent can keep narrating and handle "hold on, wrong card", and the caller never feels handed off. It is also the most work, because the suppression must apply to the model's audio or transcript path, and that is a stream your telephony vendor's DTMF masking product was not designed to know about. - Transfer to a deterministic payment module, take the call back. The agent transfers to a secure IVR or payment application, which collects, tokenises, authorises, and returns a result code; the agent resumes with
{status, token, last4}. Simplest to certify and the easiest to reason about, because the payment path has no model in it at all. The cost is a handoff in the middle of a working conversation, all the state-loss problems of escalation and warm transfer, and a caller who fails in the module has landed somewhere the agent cannot help them. - Out-of-band link, agent waits on a webhook. Send an SMS or push a link into the app, the caller pays on a hosted page, the agent waits and confirms. Best descoping available — nothing payment-related ever enters the voice path — and it fails for callers without a smartphone, adds a minute of dead air the agent has to fill, and asks a customer to follow a link sent during an inbound call, which is the exact shape of the fraud your own risk team warns about. Use it as a first choice for scheduled or follow-up payments and as a fallback, not as the primary path on an inbound call.
What none of the three do is let the model hold the digits "just in memory, just for a second". There is no such state. The context window is a request body, it is logged somewhere by default, and the compliance question is not whether you meant to retain it.
Design-review question that settles which shape you have: name the exact component that first sees the digits, then trace every stream leaving it. If one of those streams reaches the ASR, the model, the orchestrator or the tracer, you do not have a suppression architecture — you have a masking feature applied to the recording, which was the least important of the six copies.
The tool signature is the compliance artefact.
This is the one concrete thing to change today, and it is checkable in a code review rather than in an audit. A tool the model can call with a card number is a tool whose arguments will appear in your trace store, your prompt logs, your replayed eval fixtures and, eventually, in a screenshot in a support ticket. Make the data unrepresentable in the model's vocabulary.
# wrong — PAN and CVV are now model-visible and trace-visible charge_card(pan: str, expiry: str, cvv: str, amount: int) # right — the model can only ask for collection and spend a token start_card_collection(amount: int, currency: str) -> { session_id, prompt_played: true } get_collection_status(session_id) -> { state: "waiting" | "captured" | "failed", fields_done: ["pan", "expiry"], last4: "4242", brand: "visa", failure: "luhn" | "timeout" | "caller_abandoned" } authorize(session_id, amount, currency, idempotency_key) -> { status, token, auth_code, decline_reason } # the model never receives a PAN, CVV, expiry or full track data. # everything it can say about the card is last4 + brand.
Three properties of that interface are load-bearing. The model drives the flow and never the data, so all its reasoning — retry, wrong card, read the decline reason back kindly — happens over tokens and status codes. The status call is a poll over field completion rather than a stream of digits, so the agent can narrate progress ("I've got the long number, now the expiry") without ever having had the number. And authorize takes an idempotency key, because a voice agent that loses the call after sending an authorisation and retries on resume is the ordinary double-charge bug with a model in the middle; the reasoning is the same as in idempotency and retries.
Then add the negative test. A unit test that calls your tool layer with a PAN-shaped string in every argument and asserts it is rejected before serialisation, plus a CI check that greps your recorded traces and eval fixtures for Luhn-valid digit strings. The second one finds the leak the first one was supposed to prevent, and it is the single highest-value check on this page.
Callers will say the number out loud, and that is the real hazard.
Everything above handles the keypad. The failure mode that actually bites is that a caller who has been asked to type their card will read it aloud instead — because they always have, because a human agent used to accept it, and because your agent sounds like a person. Treat it as certain rather than as an edge case: some fraction of callers in every voice-first flow starts speaking digits, and the moment they do, every control you built at the DTMF layer is bypassed by a path you did not instrument.
The mitigation has to sit as early in the pipeline as you can put it, and the ordering is the whole point:
- Suppress before recognition, not after. A redactor on the finished transcript means the digits existed in a buffer, were sent to a speech vendor, and are probably in that vendor's request log. If your ASR runs in your own boundary, a digit-sequence detector on the partial hypothesis can stop the stream. If it runs at a provider, you need the detector on the audio side — energy-and-cadence detection of a spoken digit run is crude but it is in the right place.
- On detection, fail closed: drop the turn. Do not transcribe it, do not send it to the model, do not log it. Pause the recording buffer, discard the span, and have the agent re-prompt with a deterministic, non-model-generated line: "I'm not able to take card details by voice — please type them on your keypad." Constant string, not a prompt instruction, for the reason given in disclosing the agent on a call.
- Treat the detector as a control, not as the control. It has false negatives, every false negative is a potential compliance event, and a caller reading digits with pauses and corrections is genuinely hard to catch on a partial hypothesis. It reduces incidence; the architecture in Step 2 is what makes the incidence survivable.
- Pre-empt it in the prompt design. The agent should say "using your keypad" before the caller has a chance to start reading, and should never ask an open question in that turn. The cheapest reduction in spoken-PAN events is a sentence that does not invite one.
- Scan every sink continuously, not just at launch. A Luhn-valid-string detector running over transcripts, traces, eval sets and CRM notes, on a schedule. It will find something within a month of go-live, and finding it is the point; this is the sink-side discipline from PII redaction in agent traces.
Note the asymmetry that makes this harder than the human case, and worth stating plainly to whoever signs off the design: a human agent who hears a card number can be trained to not write it down, and the recording can be stopped by a button they press. A model cannot be instructed to un-receive its input, and the copies are made by the infrastructure rather than by a person's choice. "The agent was told not to" is not a control here — it fails the tamper-proof and mediation tests in the reference monitor.
Keep the callers the keypad excludes.
A keypad-only payment path is a clean compliance story and it silently drops a specific set of people: callers on a rotary or poorly-behaved handset, anyone on a relay service, callers with motor or vision impairments for whom locating twelve small keys under time pressure is the hard part, speakerphone users in a car, and anyone whose device sends DTMF the far end never receives. These are the same cohorts that accessible voice agents shows failing hardest everywhere else in the call, and here the failure is terminal — they cannot pay.
- Always have a human path for payment, and keep DTMF masking on it. This is the unglamorous consequence: automating the voice channel does not let you retire the contact centre's masking infrastructure, because the accessibility fallback runs through it. Budget for both.
- Offer the out-of-band link as an explicit alternative, not as a failure state. "I can text you a secure link instead" offered after one failed keypad attempt reads as helpful; offered after three reads as a dead end. And say the merchant name in the same breath, because you are asking someone to trust a link during an inbound call.
- Never retry the same modality three times. Two keypad attempts, then change the shape: link, callback, or human. The third identical attempt is where abandonment concentrates, and unlike a failed order lookup, an abandoned payment is lost revenue with a support call attached.
- Handle the stored-card case first, because it removes the problem. For a returning customer, a token-on-file plus a step-up authentication bound to the action — not to the call — skips this entire page. That binding is the argument in caller authentication for voice agents, and the payment step is where it pays for itself.
Operate it: five numbers and the one review that finds real defects.
Payment is the one step in the call where the agent's conversational metrics tell you almost nothing. Track it as its own funnel, and instrument the compliance surface as carefully as the conversion one.
- Payment-step completion rate, segmented by modality and by handset class. Keypad, link, transfer, human. The spread between mobile and landline is usually wide enough to change which path you offer first.
- Spoken-PAN detection events per thousand calls. Watch the direction, and treat zero as a broken detector rather than a clean channel. A step change after a prompt edit is the fastest signal you have that a wording change invited people to read the number out.
- Time from first payment prompt to authorisation. Long tails here are people struggling with the keypad, and this is where the accessibility failures from Step 5 show up as a latency distribution before they show up as a complaint.
- PAN-shaped strings found per sink per scan. Transcripts, traces, prompt logs, eval fixtures, CRM notes. The target is zero and the value of the metric is entirely in the non-zero months.
- Declines the agent handled without a human. A decline is the most common outcome you will under-design for, and the agent needs a deterministic script per decline reason — reading a raw processor message to a caller is both unhelpful and occasionally a disclosure.
Then run one review that no dashboard replaces: take ten recorded payment steps and trace each one through every artefact your stack produced, by hand, opening each store. Not a policy review — a search. You are looking for the copy nobody designed, and in a stack assembled from a telephony vendor, a speech vendor, a model provider, an orchestrator and a tracer, there is usually exactly one.
If you do one thing: change the tool signature. Replace any tool that can accept a card number with start_card_collection / get_collection_status / authorize(session_id, …) so that a PAN is not expressible in anything the model can emit, then grep your existing traces for Luhn-valid digit strings to find out what the old signature already cost you. Everything else on this page is architecture with a lead time; that pair is an afternoon, and it closes the copies that are hardest to find later. Related: tool use and state in voice for the surrounding turn design, shopping and checkout agents for the same payment step in a text channel, and agent payments for where tokenised, agent-initiated payment is heading.