Payroll & Employment-Tax Agents

11 min read

Y40
Playbook · Domain Playbooks

Payroll and employment-tax agents.

Payroll is the one business process that already has a deterministic authority in it. A rules engine loaded with rate tables computes the number, and that number is what gets deposited, filed and printed on a pay statement — so an agent that produces or edits a figure destined for a filing has inserted a probabilistic step into a process whose entire legal and penalty structure assumes a determinate one. Build it the other way round and the same agent is worth a headcount: let it read everything, write only into inputs and drafts, and spend its whole budget on the reconciliation, the registration gap and the agency notice that a human would otherwise find in November. The constraint that should shape your design is not accuracy, it is the calendar — payroll cannot be late, so an agent that responds to uncertainty by asking a question and waiting has produced a compliance failure with a penalty attached.

STEP 1

Draw the line at the engine, and make it architectural rather than a policy.

Every payroll platform contains a calculation engine whose outputs are auditable, reproducible and — this is the part that matters — defensible. When a state agency disputes a withholding amount, the answer is the engine's rate table version and its itemised computation, not a rationale. Your agent must be on the input side of that engine and on the reading side of its output, and never in between.

# the only three write paths an agent should have
INPUT     propose a change to source data (address, exemption,
          deduction, work location) -> staged, human-confirmed
DRAFT     assemble a return, a notice reply, an amendment packet
          -> unsigned draft in the filing queue
NOTE      annotate a case, a variance, a reconciliation item
          -> free text, never read back as a number

# everything else is read-only, including:
the calculated number   the deposit amount   the filed return
the pay statement       the GL journal       the agency portal

The reason this has to be enforced in the schema rather than in the prompt is the asymmetry of the failure. An agent that proposes a wrong input gets caught by the confirmation step and by the gross-to-net variance check in Step 3. An agent permitted to write a number directly produces a deposit that is wrong by an amount nobody can reconstruct, and the reconstruction is what the penalty assessment will ask for. This is the same argument as consumer credit agents keeping scoring in a declared model the agent cannot reach, and it generalises: wherever a deterministic authority already exists, the agent's job is to feed it and to audit it.

One concrete test for whether you have drawn the line correctly: ask what the agent does when the engine's output looks wrong to it. The correct behaviour is to raise a variance with evidence and stop — not to adjust an input until the output matches its expectation. An agent that can iterate on inputs until the number looks right has been given the ability to produce any number at all, through a path your audit trail records as a series of legitimate data corrections.

STEP 2

The primary control is a timer, not a confidence threshold.

In most domains, an agent that is unsure and escalates has behaved well. In payroll, waiting is itself the error, because the obligations are dated and the penalty is a function of lateness rather than of magnitude. US federal deposit penalties step from 2% at one to five calendar days late, to 5% at six to fifteen days, to 10% beyond fifteen, to 15% once a notice for immediate payment has gone unanswered for ten days — a ladder that replaces the lower rate rather than adding to it, and that is indifferent to whether the delay came from a hard question or an unread message.

  • Every agent task carries a deadline derived from the obligation, not from a queue policy. A semiweekly depositor's obligation, a quarterly return, a 31 January statement deadline, a state's own new-hire reporting window. The task's SLA is the obligation's date minus your internal review time, and it is set when the task is created.
  • "I don't know" routes; it never blocks. Uncertainty produces an escalation to a named person with the deadline attached and a default action stated. An unclaimed escalation re-routes upward on a timer. The pattern is the review queue in review queues for agent output, with the addition that the queue has a clock and the clock has consequences.
  • Work the calendar backwards, and start early enough to be wrong once. The agent should open the registration gap, the missing tax form and the unreconciled variance early in the period, because the remedy for most of them takes days of somebody else's time. An agent that surfaces a missing state registration on the deposit deadline has technically succeeded and practically failed.
  • Scheduled runs need their own failure alarm. A payroll agent that did not run is indistinguishable from a payroll agent that found nothing, which is the fail-open shape in its most expensive form. Assert on the heartbeat, per scheduled and triggered agents, and treat a silent period as an incident rather than a quiet week.

Worth stating plainly because it inverts the usual instinct: in this domain you should prefer an agent that acts on a documented default and flags it to one that pauses for a human. A default deposit made on time and corrected next period costs interest; a missed deposit costs a tiered penalty, a notice, and an hour of a controller's attention for each of the next three months.

STEP 3

Spend the capability on reconciliation and exceptions, where the ground truth already exists.

The work that actually consumes a payroll team is not calculation. It is a long tail of comparisons between systems that should agree and do not, each requiring someone to read two records and a rate table and decide which one is wrong. That is a near-ideal agent task: high volume, low creativity, and a deterministic check available for every answer.

  • Period-over-period variance on gross-to-net. Per employee, per earning code, per tax. Flag a change without a corresponding input change — a new deduction, a rate-table update, a jurisdiction change — and name the candidate cause. This one check catches most of what would otherwise become a W-2 correction.
  • Payroll-to-general-ledger reconciliation. The recurring, tedious, deferrable task that is also where an unnoticed misposting compounds for a quarter. Same shape as the matching work in accounts payable agents, and it wants the same treatment: propose the journal correction, never post it.
  • Agency notice triage. Notices arrive as scanned PDFs, in dozens of formats, from dozens of agencies, and most are informational. Classify, extract the deadline and amount, join to the period, and draft the reply with the supporting register attached. Extraction accuracy should be measured on the population that matters — the notices that carry a deadline — and not on the whole mailbox, which is the error consumer credit agents makes in its own tail.
  • Return-to-register reconciliation before filing, not after. Quarterly wages reported versus the sum of the period registers; annual statements versus the four quarterly returns. Discrepancies found before submission are edits; found after, they are amendments.
  • Registration and account coverage. For every work location present in this period's data, does an active employer account exist for every tax that location triggers? This is the check that catches remote-work drift, and it is pure lookup — the agent's contribution is that it runs every period instead of once a year.

Give each of these a deterministic verifier and report the agent's precision on the verifier rather than a model-graded score. Variance detection is checkable against the input diff; GL reconciliation against the trial balance; notice extraction against the eventual human correction. A domain with this much available ground truth is one where an LLM-as-judge metric is a choice to measure something weaker than what you can measure for free.

STEP 4

Never let the model infer a jurisdiction; make it call the lookup and stop on a miss.

The United States has more than 7,400 local taxing jurisdictions and adds roughly 200 a year. Pennsylvania alone has over 2,500 municipalities and nearly 500 school districts levying earned income tax under Act 32; Ohio has 649 municipalities and 199 school districts with income taxes, plus joint districts that extend a city's tax past the city limits. No model holds this correctly, the ones that hold it approximately are the dangerous case, and the boundaries do not follow postal addresses.

  • Address to jurisdiction code is a tool call, and its failure is terminal. The agent may not reason about which municipality a street is in, may not pick the nearest match, and may not fall back to the state default. An unresolved address becomes an exception with a deadline — see the terminal-versus-retryable distinction in retry amplification, which applies exactly here.
  • Residence and work location are two answers, and both are needed. Reciprocity agreements, resident credits and local splits all depend on the pair. An agent that collects one and assumes the other has produced a plausible wrong withholding, which is the worst kind because it survives review.
  • A new work location is a registration workflow, not a data field. One employee moving states can trigger an income-tax withholding account, an unemployment-insurance account, new-hire reporting, a local account, and sometimes a paid-leave programme. The agent's job is to open that checklist the day the address changes and to chase it, which is the casework pattern applied inward.
  • Rate-table versions belong in the record. Pin the table version used for every calculation the agent reasoned about, and log it next to the decision. When an agency disputes an amount two years later, the question is which table was in effect, and audit trails that record the conclusion without the table version cannot answer it.

The general rule this instance illustrates is worth keeping: in any domain with a published authoritative table, the agent's value is in knowing that a lookup is required and in noticing when one is missing — never in remembering the contents. Hallucinated tax rates are rare; hallucinated jurisdictions are not, because an address genuinely looks like enough information to reason from.

STEP 5

The employee-facing surface is the highest-volume use and the easiest to get wrong.

"Why is my paycheck different this month" is the most common payroll question in every company, and answering it well is worth more to employees than anything else on this page. It is also where a correct paycheck and an incorrect explanation combine into a real harm: the employee changes a withholding election, disputes a deduction, or loses trust in a number that was right.

  • Explain from the engine's itemisation, never by recomputing. The answer is a diff of two itemised statements plus the input change that caused it. If the agent cannot point to an input change, the answer is "I don't know yet, here is who is looking", not an arithmetic reconstruction that happens to land close.
  • Route, do not answer, on the categories where being helpful is harmful. Garnishments and court orders, immigration status, leave of absence and disability pay, termination and final-pay rules, anything touching a benefits election deadline. Each has a specialist and a legal exposure, and a confident paraphrase is the failure mode. This is the address-and-authority split from human in the loop, applied by topic rather than by risk score.
  • Treat the pay statement as a document with legal force. In several jurisdictions the itemised statement is a mandated disclosure with its own penalties, so an agent-generated summary presented alongside it must be visibly a summary and must not restate any figure it did not read from the statement.
  • Never ask for what you already hold, and never echo more than you were asked. A payroll agent has access to compensation, garnishments, benefit elections and dependants. The retrieval scope for an employee's own question is that employee, enforced by the query and not by the prompt.

The metric to watch here is not satisfaction, it is the rate at which an agent explanation is followed by an employee-initiated change that is later reversed. That number isolates the specific harm — a confident explanation that caused a wrong action — and it is measurable from data you already keep. See production feedback signals for the general shape.

STEP 6

Amendments cascade, so treat the first filing as the only cheap one.

The economics of a payroll error are asymmetric in a way that should change how much review you buy. Correcting a quarterly return is a separate amended filing; correcting a wage statement means a corrected statement, a transmittal, and — if the year has closed — a cascade into an employee's own personal tax return, which you cannot fix and they must. Each one consumes specialist time, and the specialist is the person your agent was supposed to free up.

  • Separate prepare from submit, permanently. The agent assembles; a credentialed human signs. A signature on a return is a personal legal attestation, and no authorisation from your own organisation changes that — which makes this the one approval in the system that cannot be delegated to a policy engine. The binding mechanics are in approval and confirmation UX.
  • Weight review by irreversibility, not by amount. A large payment that can be corrected next period is cheaper than a small wage-statement error discovered in February. Sort the review queue by how expensive the correction path is, which for most payroll objects means: statements first, returns second, deposits third.
  • Keep the agent's evidence longer than you keep its traces. An amendment argued two years later needs the register, the rate-table version and the input diff the agent used — a small, structured record that should outlive the full trajectory under a policy set deliberately, per retention and legal hold.
  • Measure the agent on corrections avoided, not on tasks completed. Amendments per thousand employees per quarter, before and after. It is the only number that distinguishes an agent that is doing the reconciliation work from one that is generating plausible reconciliation notes, and it is the number a payroll manager already tracks.

Start with agency notice triage and period-over-period variance, in that order, and ship both read-only for a full quarter before you let either propose an input change. Notice triage pays for itself immediately because the alternative is a mailbox nobody owns, and variance detection is where you will discover that your ground truth is better than you thought — most payroll systems can tell you exactly which input moved, and nobody has ever asked them per employee per period. Do not begin with the calculation, do not begin with filing, and do not let the first version have a write path to a number: the ceiling on this agent's value is set by how much of the exception tail it can work through before a deadline, and none of that requires it to be the authority on anything.