Smart-home agents: the actuation is the demo, the event history is the product and the liability.
Every smart-home agent demo turns on a light, and every smart-home agent incident will be about something it read. Device event history is a presence log for everyone who lives in the building — attacker-writable at the device-name layer, unbounded in the time dimension, and impossible to un-read once it is in a context window — while the actuation half is the part that is enumerable, reversible and easy to gate. Build the read path with a budget and a window cap before you write a single command, and treat "turn on the lamp" as the easy problem it is.
Sort devices by reversibility, not by category, and grant autonomy per class.
The useful axis is not lights-versus-locks. It is what happens when the agent is wrong and nobody is watching. Three classes, and they earn different permissions:
- Reversible and low-consequence. Lamps, scenes, media, a speaker's volume. A wrong action is an annoyance, corrected by the opposite command. This is the only class that earns unattended autonomy, and it is most of what people actually want.
- Reversible but consequential. Thermostats, blinds, irrigation, EV charging. A wrong action costs money, comfort or a frozen pipe before anyone notices, because the feedback loop is hours long. Autonomy here needs bounds — a temperature range, a duty cycle, a daily budget — not a confirmation prompt.
- Irreversible or safety-bearing. Locks, garage doors, gas and water valves, ovens, anything medical, anything that admits a person to the building. No autonomy. Not behind a confirmation, either: a confirmation that fires often enough becomes a reflex, and the one that matters arrives looking like the other forty.
Platforms are starting to enforce the third class for you — Google's Home MCP blocks door unlocking outright rather than exposing it behind approval — and that is the right shape to copy. Refusing to expose a capability is a stronger control than gating it, because it cannot be worn down by frequency. Where you must expose one, require a second factor that is not the agent: a physical press, a separate app confirmation, a code the agent does not hold. See physical actuation safety for the general treatment.
Write the class onto the device record, not into the prompt. A prompt-level rule is advisory and gets summarised away over a long session; a class attached to the device is checked by your tool wrapper on every call, including the call at step forty that nobody reviewed.
Your action space is discovered at runtime, so policy has to be written over traits.
This is the part that breaks a conventional allowlist. Home platforms expose devices through capability traits, and the commands available on a given device come from a schema the platform reports when you connect. Someone pairs a Matter blind on a Tuesday and the agent's action set has grown — with no deploy, no review and no version bump anywhere in your stack.
- Enumerate at connect and on change, then classify. Pull the resource list, map every trait to one of the three classes in STEP 1, and persist the mapping with a version. The mapping is your policy artefact; the device list is not.
- Default-deny unknown traits. A trait your classifier has never seen is not low-risk by default. It is unclassified, which means unavailable until someone classifies it. This single rule is what keeps a new product category from silently entering the autonomous tier.
- Diff the surface and alert on growth. "Three new traits appeared on this account" is an event worth a notification, exactly like a new IAM permission. Tool-catalogue lifecycle is the general discipline; here the catalogue changes because someone went shopping.
- Never let the model choose the trait mapping. Asking the model "is this action safe" at runtime puts the classification inside the thing being constrained. Classification is a build-time decision with a human in it.
The physical world has no transactions. Build a read-modify-verify contract.
Nothing in a home rolls back, and a command that times out is in an unknown state rather than a failed one. Most agent stacks treat a timeout as a failure and retry, which is how a garage door ends up cycling.
- Express commands as absolute targets, never as relative deltas.
set_temperature(20)is idempotent;turn_up_by(2)applied twice is a different house. This single rule removes most retry damage — the same argument as idempotency and retries, with physical consequences. - Verify with an observation, not with the command's return value. A 200 means the platform accepted the request. Whether the device did it is a separate question answered by reading state back, with a bounded wait.
- On timeout, read before you retry. Never re-issue blind. If the state read is also unavailable, surface an explicit unknown to the user and stop — an honest "I could not confirm the garage door closed" is worth more than a confident summary.
- Sequence, do not fan out, within a scene. A batch that half-applies leaves the home in a state no one designed. Apply in a defined order, verify each step, and report the partial state plainly if you stop midway.
// Absolute target, explicit verification, unknown is a first-class outcome.
{ "action": "set_temperature", "device": "thermostat.hall", "target_c": 20 }
→ accepted
{ "verify": "read_state", "device": "thermostat.hall", "within_ms": 8000 }
→ { "setpoint_c": 20, "observed_at": "2026-09-19T18:04:11Z" } // confirmed
// Timeout path — do NOT re-issue the command.
{ "verify": "read_state", "device": "garage.main", "within_ms": 8000 }
→ { "status": "unknown", "reason": "device unreachable" } // tell the user, stop
Event history is the sensitive asset. Budget it like a cost, not like a lookup.
A home platform's history API returns past state changes and event logs over whatever window you ask for. Each row is innocuous — a light on, motion detected, a door opened. The set is a behavioural record: when the house is empty, who comes home when, which nights nobody slept there. No allowlist can refuse that, because the sensitive object is the pattern, and the pattern is an inference the agent draws.
- Cap the window server-side, in your own wrapper. Default to hours. Make a wider query a separate, logged, user-visible request. A width limit is the one control that survives every downstream mistake, and it is a few lines.
- Fetch for a stated purpose and discard. "Why was the heating on all day Tuesday" needs Tuesday, not the month. Pull the narrow window, answer, and drop it rather than letting it ride in context for the rest of the session.
- Keep it out of memory and out of summaries. A persistent-memory write derived from event history outlives the question that justified it, unattributed. Exclude the history tool's output from memory extraction explicitly — this is one of the few cases where a blanket exclusion is right.
- Know your retention chain. Once history enters a context window it is in your traces, your provider's logs, and any debugging export. Decide whether those systems should hold it, and redact at the boundary if not — see PII redaction in agent traces.
- Prefer aggregates when the question is a pattern. If the user wants "am I leaving the heating on too much", compute it and return the number. Do not hand the raw log to the model and ask it to notice.
Treat every string from the home as untrusted input, starting with the device names.
The realistic compromise is not a hostile agent. It is a correct agent reading text somebody else wrote, and a home is unusually good at producing that text.
- Device and room names are attacker-supplied. Anyone who can add or rename a device on the account controls a string that lands in the agent's context — a guest, a contractor, a previous occupant, a compromised third-party integration. Render names as data, never as instructions: keep them in a delimited field, strip control characters, cap length, and never interpolate one into a system-level instruction.
- Camera and doorbell event summaries are model output. A description of what a camera saw is generated text, produced by a model that has no idea its output will be read as fact by something holding actuation tools. That is a two-model supply chain; treat the upstream one as an untrusted source. Telemetry as untrusted input is the pattern.
- Assume injection plus reconnaissance arrive together. The connection that could carry an injected instruction also carries the household's schedule and the command surface. Design so that a successful injection cannot reach the class-three devices at all — which it cannot, if STEP 1's classes are enforced in your wrapper rather than in the prompt.
- Log the provenance of anything that influenced an action. When an agent acts on an event, record which event, from which device, at what timestamp. Reconstructing "why did the lights come on at 3am" six weeks later is the ordinary case — see audit trails.
A home is multi-person; an OAuth grant is not. Design for the people who never saw the consent screen.
This is the requirement most teams discover late, and it has no purely technical fix. One account holder authorises the integration. Everyone else in the building — partners, children, housemates, guests, carers, a cleaner — is recorded by the sensors and represented in the history, and none of them consented to a third-party agent reading it.
- Make the grant narrower than the account. Share the minimum set of devices into whatever identity the agent uses, and keep cameras out of it unless a specific feature needs them. The unit of blast radius is the grant, so make the grant small — blast radius, applied to a building.
- Record what was granted, by whom, when, and how it is revoked. A single revoke path that actually works, tested, is worth more than a granular permissions UI nobody opens. Delegated access and consent records covers the artefact.
- Do not let the agent answer questions about people. "Was my daughter home last night" is technically a history query and substantively surveillance. Decide the policy deliberately, implement it as a refusal on the query shape rather than a hope about the model, and give the refusal a reason.
- Disclose in the home, not only in the app. If an agent can act or observe, the people living there should be able to find that out without having the account. A visible state — a light, a card on a display, a monthly summary to the household — is the closest thing to consent the architecture allows.
- Plan for the household changing. People move out. An agent with standing access and a memory of the previous occupant's routines is a problem with a name; make account membership changes trigger a memory and history purge.
Build the first version with the actuation tool removed entirely. An agent that can only read — "why is the house cold", "did the back door get left open", "what is costing me money" — is genuinely useful, ships in a week, and forces you to solve the window caps, the untrusted-name handling and the memory exclusions while the stakes are low. Then add class-one devices, and only then argue about the rest. The teams that do it in the other order ship the lamp demo and discover the history problem in a support ticket. Related: physical actuation safety for the write path in depth, MCP security anti-patterns for the integration layer, and undo and reversibility for what to offer when there is no undo.