Disclosing the agent on a call.
The one sentence your voice agent is legally required to say is also the single sentence in the call most likely to be talked over — you built barge-in on purpose, the caller starts speaking at 400ms, and your VAD cancels the disclosure mid-word with no record that it ever played. Treating disclosure as a line at the top of the script is the mistake; it is a state machine with re-trigger conditions, an acknowledgement requirement, and an audit artifact, and the part that actually decides your liability is what happens when someone interrupts to ask "wait, am I talking to a robot?"
Three different obligations, and only one of them is the opener.
Teams build one disclosure because they read one law. The rules in force in 2026 are shaped differently from each other, and a script that satisfies the loudest one still fails the others.
- Up-front, unprompted. The EU AI Act's Article 50(1) requires that a person interacting with an AI system be made aware of it, in a clear and distinguishable manner, at the latest at the time of the first interaction. Those obligations became applicable on 2 August 2026, with the Commission's final implementation guidelines adopted on 20 July 2026. California's SB 1001 sits in the same family for bots used to incentivise a commercial transaction or influence a vote.
- On request. Utah's AI Policy Act, as amended by SB 226, turns on a clear and unambiguous request: if the person asks whether they are dealing with AI, you must say so. Nothing triggers until they ask — and then the obligation is immediate and unconditional.
- Prominent and repeated, for high-risk interactions. Utah additionally requires regulated occupations to disclose prominently and up front in high-risk contexts — health, financial, legal, mental-health advice, sensitive data. Its mental-health-chatbot rules go further still: disclose before the user can access any feature, again at the start of any interaction where the user has not used it in the previous seven days, and again any time the user asks.
The federal telephony layer is a separate axis and is about consent, not disclosure. The FCC's February 2024 declaratory ruling put AI-generated voices inside the TCPA's "artificial or prerecorded voice" definition, which means prior express consent before the phone rings and an identification of the calling business at the start of the call. A follow-on Notice of Proposed Rulemaking from August 2024 would have added explicit AI disclosure requirements; as of 2026 it has not been finalised. Build to the ruling, not to the proposal — but note the proposal describes the direction, so do not architect anything that could not add a disclosure obligation later.
The design consequence of the three shapes together: you need an unprompted disclosure, an on-demand disclosure, and a re-disclosure rule, and they are three separate code paths. Most stacks ship the first, fake the second with a prompt instruction, and have never heard of the third.
Barge-in is actively hostile to your opener. Design around it.
Here is the failure nobody tests for. Mandatory barge-in — see turn-taking and barge-in — means the moment the caller produces speech energy, you stop talking. Callers begin speaking early on answered calls, particularly on inbound support lines where they have been waiting and arrive with a prepared sentence. Your disclosure is the first thing out, so it is the utterance with the highest probability of being cut.
Worse, the cut is silent. A TTS stream that is cancelled at 900ms into a 2.4s utterance logs as "played." Your compliance evidence says the disclosure happened. The caller heard "You're speaking with an auto—".
Four fixes, in the order they are worth doing:
- Make the disclosure short enough to survive. Target under 1.5 seconds. "This is an automated assistant from Northgate Energy." Every clause you add is another chance to be interrupted, and none of the regulations require a paragraph.
- Put it first, before any greeting or menu. Not after "Thanks for calling!" — the greeting is exactly what invites the caller to start talking.
- Suppress barge-in for the disclosure specifically. A single non-interruptible window at the very top of the call. This is a deliberate exception to a rule you otherwise want universally, and it is defensible precisely because it is bounded to one short utterance.
- Log completion, not dispatch. Record the disclosure as satisfied only when the TTS stream reached its final frame. If it was cut, set a flag that the re-disclosure rule in Step 3 will pick up.
If you suppress barge-in, do it for under two seconds and never for the menu that follows. A non-interruptible disclosure is a compliance control; a non-interruptible IVR tree is the thing your callers hate, and replacing an IVR exists because of it.
Every entry path into the agent needs its own disclosure.
"At the latest at the time of the first interaction" is doing more work than it looks like. The first interaction is per person and per interaction — not per call leg, and certainly not per session ID in your database. Enumerate the ways a human ends up talking to your agent without hearing the opener:
- Transfer in from a human. An agent takes over from a representative, or a warm transfer runs the other direction. The new party on the line has heard nothing.
- Callback and redial. Your system rings back. The disclosure state from the original call does not carry.
- A second human joins. Conference, speakerphone handoff, "let me put my wife on." The person who just arrived is a first interaction.
- Resumed session after a drop. The caller redials within thirty seconds and you restore context. Restoring context is not the same as restoring consent — and note Utah's seven-day rule for its mental-health case, which is an explicit legislative judgement that disclosure expires.
- IVR to agent. The caller spent ninety seconds in a deterministic menu and is then handed to a model. That is the moment the AI interaction starts.
- The cut opener from Step 2. Flagged, and never re-delivered by most implementations.
Implement this as an explicit predicate rather than as script placement, because script placement cannot see any of the six cases above:
# Disclosure is a property of the (human, interaction) pair. must_disclose(party) := not party.heard_completed_disclosure or party.joined_after_last_disclosure or transfer_occurred_since(party.last_disclosure) or now - party.last_disclosure > RE_DISCLOSE_WINDOW or party.asked_if_ai # always, unconditionally # Evaluate on: call start, transfer, party join, session resume, # and on every turn where the intent classifier fires "is_this_a_bot". # Emit a disclosure_delivered event with: party id, utterance text, # completed_to_final_frame, timestamp, trigger reason.
That event is the artifact. When a regulator or a plaintiff asks whether a particular caller was told, the answer is a row, not a screenshot of a prompt template. This is the same reasoning as audit trails: a control you cannot evidence per-interaction is a control you cannot demonstrate at all.
"Are you a real person?" is an intent, not a prompt instruction.
This is the part that separates a compliant system from one that merely looks compliant, and almost nobody builds it.
The Utah obligation, and the practical core of every other regime, triggers on the caller asking. If your handling of that question is a line in the system prompt — if the user asks whether you are an AI, tell them you are — then your compliance with a statute is a model behaviour, subject to everything model behaviour is subject to: a long context that buries the instruction, a persona prompt pulling the other way, a jailbreak-shaped customer, a model swap during a fallback. It will hold ninety-something percent of the time, which is a fine number for tone and an unacceptable one for a legal obligation.
Build it as a deterministic path:
- A dedicated intent classifier on the transcript stream, running on every user turn, firing on the family: are you a real person, am I talking to a bot, is this a recording, are you AI, are you human, can I speak to a person. Cheap, and it is one of the few intents where a false positive costs you nothing.
- A fixed response, not a generated one. The answer is a constant string. It does not get paraphrased, softened, or wrapped in personality. This also removes any chance of the model hedging its way into a false denial, which is the specific outcome that turns a compliance question into a deception question.
- Priority over whatever the agent was doing. The question interrupts the task. Answer it, log it, then offer to continue.
- A human-escalation offer attached. Many callers asking the question are really asking for a person. Coupling the two makes the honest answer feel like service rather than a brush-off.
Then test it adversarially, because the interesting cases are not the plain ones: the caller who asks in the middle of a barge-in, who asks in a second language (multilingual agents need the classifier in every language you answer in), who asks obliquely — "you sound like a computer" — and who asks at turn forty when the context is full. Put those in the eval set with a hard pass threshold of 100%, and treat a miss as a release blocker rather than a quality regression.
One rule worth writing into the system prompt anyway, as defence in depth behind the deterministic path: the agent must never affirmatively deny being an AI, never claim a human name as its own identity, and never say "I'm a person." A model that hedges — "I'm here to help you personally!" — has produced a deception, and the difference between a disclosure failure and a deception finding is the difference between a fine and a headline.
Write the sentence once, carefully, and stop workshopping it.
Marketing will want the disclosure warm. Legal will want it complete. Both instincts make it longer, and Step 2 already established that length is the enemy. The constraints that actually bind:
- Name the automation plainly. "Automated assistant", "AI assistant", "virtual assistant". Avoid a bare product name — "You're speaking with Aria" discloses nothing, and a human-sounding first name with no qualifier is the pattern regulators single out.
- Name the business. The TCPA identification requirement is separate from the AI disclosure, and folding them into one clause costs you nothing: "This is an automated assistant from Northgate Energy."
- Do not stack the recording notice into the same breath. Consent to record is a different legal question in a different set of jurisdictions — see recording consent and redaction — and jamming both into one utterance doubles the length of the thing you most need to survive barge-in. Disclosure first, recording notice second, both short.
- Make the on-request answer more direct than the opener. "No, I'm an AI assistant — I'm not a person. Would you like me to transfer you to someone?" Directness is the point; this is the sentence a transcript will be read back from.
- Version it and pin it. The exact string, with a version identifier, stored next to the disclosure event. When the wording changes you need to know which callers heard which text, and an A/B test on this string is a compliance change, not a copy change.
Translate it properly for every language you answer in, and have it reviewed rather than machine-translated. A disclosure that is unclear in the caller's language does not satisfy a requirement written as "clear and distinguishable."
What to ship, and what it costs you when you skip it.
The honest economics: disclosure is widely assumed to suppress conversion, and the assumption drives a lot of quiet non-compliance. Two things are worth saying about it. First, on inbound support lines the effect is small and sometimes positive, because callers who know they are talking to a machine phrase requests more literally and the agent's recognition improves — the disclosure is doing work for you. Second, on outbound the effect is real, and it is also the channel with the sharpest regulatory exposure, so the cost is the price of operating there rather than an argument against disclosing. See outbound voice agents for what else that channel demands.
The checklist, in build order:
- A single short disclosure utterance, non-interruptible, first thing on the call, logged only on completion to final frame.
- A
must_disclose(party)predicate evaluated at call start, transfer, party join, session resume, and long-call boundary — not a line in a script. - A deterministic "are you a bot" intent path with a constant-string answer, priority over the current task, and a transfer offer attached.
- A hard prohibition on affirmative denial, in the prompt and in the eval.
- A
disclosure_deliveredevent with party, text, version, completion flag, timestamp and trigger — retained as long as the call recording. - Eval cases at 100% pass: cut opener, transfer-in, second party joins, asked mid-barge-in, asked in each supported language, asked at turn forty.
Run this audit tomorrow and you will know where you stand in an hour. Take last week's calls, filter to those where the transcript contains any form of "are you a real person / is this a bot / am I talking to a machine", and read every single one. You are looking for three things: how many there were (usually far more than the team guesses), how many got a clear yes, and how many got a hedge. Then filter separately for calls that began with a transfer and check whether the disclosure ever played. Those two queries, run against traffic you already have, will tell you more about your exposure than any policy review — and if the second query returns nothing because you never logged the event, that is the finding.
Related: caller authentication for the other identity question on the call, the EU AI Act for agents for the obligations around this one, and disclosure and content provenance for the same requirement in text and media.