The principal–agent problem.
Ask "who does this agent work for" and you will find at least three answers, all true at once, and the one that wins a tie was decided by someone who is not in the room. That is the principal–agent problem, and the AI version is sharper than the human one for a reason nobody expects: the agent has no interests of its own to trade with, so it does not negotiate between its principals — it silently obeys whichever one the architecture privileged. Which means the conflict is settled in your system prompt, your tool list and your vendor's post-training, not in the moment a user notices something is off.
The term is borrowed, and every part of it transfers except one.
In economics, a principal–agent problem exists when one party acts on another's behalf, their interests are not identical, and the principal cannot fully observe what the agent did or why. A real-estate agent sells your house faster and cheaper than you would, because their commission on the last $20,000 is small and their time is not. A fund manager takes risk you would not take, because the upside is shared and the downside mostly is not.
Every element of that structure is present in a deployed AI agent. It acts on someone's behalf. Its objective is set by someone else. And the person it acts for cannot see the reasoning, the tools it declined to use, or the option it never surfaced — which is the observability half of the problem, arriving in a much more extreme form than a quarterly report.
One element does not transfer, and it is the one people assume is protective.
- A human agent has interests of their own. That is the source of the conflict — and also a limit on it. A broker who is asked to do something outrageous has a licence, a reputation and a conscience to weigh against the fee.
- A model has none. It will not skim, and it will also not resist. Where a human agent trades off between principals, a model resolves the conflict by whatever ordering its training and its prompt imply, with no friction and no tell.
So "the model has no ulterior motive" is true and is not reassuring. It means the conflict does not surface as misbehaviour you could catch. It surfaces as consistent, cheerful, well-written compliance with a priority list the user never saw.
Your agent has at least three principals. Rank them before something else does.
Name them explicitly, because the exercise usually surprises the team that builds the thing.
- The user — the person who typed the request and bears the consequence. Usually the one everybody says the agent serves.
- The deployer — whoever wrote the system prompt, chose the tools, set the refusal policy and pays the bill. Their interests overlap with the user's most of the time and diverge at exactly the moments that matter: cost, deflection, retention, liability, which products get mentioned.
- The model vendor — whose post-training decides what the agent does when the first two conflict, and which the deployer can influence but not override. Its priorities live in a system card if you are lucky.
There is a fourth that is not a principal at all but behaves like one: whoever wrote the text in the context window. A retrieved web page, a tool result, an email in the inbox the agent was asked to triage. It has no legitimate claim on the agent's loyalty, which is the entire reason the instruction hierarchy exists — a written ordering of whose words count as instructions. Prompt injection is precisely the case of a non-principal successfully impersonating one.
Here is the uncomfortable structural fact. The instruction hierarchy — the mechanism that settles conflicts between principals — is written and trained by one of the principals. That is not a scandal; someone has to write it, and vendors have been reasonably careful. But it is a reason to read yours rather than assume it, and a reason a deployer cannot promise a user an ordering the vendor has not implemented.
The conflicts that matter are architectural, not ethical.
Nobody sets out to build a disloyal agent. Conflicts arrive as ordinary product decisions, and the tell is always the same: an objective the user cannot see, expressed as a metric someone is measured on.
- A support agent that both helps you and protects deflection rate. Escalating to a human is the right answer for one principal and a miss for the other. Whichever one is in the eval is the one that wins, every time. See customer-support agents.
- An agent whose vendor is paid per token. Verbosity, retries and enthusiastic tool use are not corruption; they are the absence of a countervailing pressure. Nobody has to intend this for it to hold.
- A recommendation drawn from a catalogue somebody paid to be in. Affiliate content in retrieved pages, a supplier list, a sponsored placement. The agent is not lying; its evidence was selected by a party with an interest. This is the everyday case and it is covered in commercial influence and paid placement.
- An agent acting for two users at once. A scheduling agent negotiating between your calendar and mine; an agent buying on your behalf from a merchant's agent. Here the principals are symmetric and the question "whose agent is this" has a real answer that must be stated, not inferred. Agent payments is where this bites first.
- Agreeableness as a loyalty failure. Sycophancy is usually filed under quality. It is better understood here: an agent that confirms your premise is serving your comfort over your interest, and comfort is the principal that shows up in thumbs-up data.
Note what all of these share. None requires a bad actor, a jailbreak, or a model failure. Each is a correct agent optimising a stated objective, where the statement of the objective is where the loyalty was decided. That is why "we'll add a guideline about it" does not work — a guideline in the prompt competes with a metric in the eval, and the metric wins.
Make the principal legible, and stop asking one agent to hold two objectives.
Four moves, in rough order of how much they buy.
- Write the ordering down, in the system prompt, in the words a user would recognise. Not "be helpful" — "when the user's interest and ours conflict, do X". An ordering you cannot write in one sentence is an ordering you have not made, and the model will infer one anyway.
- Separate the agents rather than blending the objectives. One agent that both advises and sells is a worse design than two agents with disclosed roles, because a blended objective is unauditable — there is no counterfactual to compare the answer against. Splitting them costs a handoff and buys a check anyone can run.
- Put the conflicted step in front of a human, and say which step it was. Human-in-the-loop is expensive enough that it should be spent where the interests diverge, not sprinkled over everything. A refusal or an escalation that names its reason — see structured refusal — is the artefact that makes the conflict visible afterwards.
- Measure the conflicted metric against the user metric, in the same review. Deflection rate next to resolved-without-recontact. Conversion next to return rate. A single number can always be moved by serving the wrong principal; a pair cannot, which is the cheapest structural honesty available in evaluation.
And keep the question alive as systems compose. When your agent calls another organisation's agent, the second one's principal is not you — that is what agent identity and permissions is ultimately for, and accountability and roles is where it gets written down for a regulator.
Take your agent's system prompt and highlight every sentence that serves someone other than the person typing. Most teams find two or three, and are surprised by one of them. Then ask the harder question: if the user could read the highlighted lines, would you still ship it as written? If yes, you have a defensible design and you should consider showing them — transparency is cheap when you have nothing to hide. If no, you have found the conflict, and a guideline further down the prompt is not going to resolve it. Related: goal drift for what happens to the ordering over a long run, and the instruction hierarchy for the mechanism that enforces it.