Consumer Credit Agents

8 min read

Y39
Playbook · Domain Playbooks

Consumer credit agents: three of the four regulated acts happen before anything is decided.

Teams building an agent for consumer lending brace for the hard part — explaining the decision — and then get caught by the easy part, because in credit the regulated conduct starts the moment the agent opens its mouth. What it asks is restricted. What it says to a hesitant applicant can be unlawful discouragement. The point at which it stops collecting documents silently starts a thirty-day notice clock nobody wired a timer to. Only the fourth act, the adverse action reason, is the one everyone prepared for. Build the conversation as a regulated interview with an enumerated question set and a clock, keep the scoring in a declared model the agent cannot reach, and the explainability problem stops being the interesting one.

STEP 1

Map the four regulated acts before you map the workflow.

In the United States, Regulation B (12 CFR Part 1002, implementing ECOA) governs the whole interaction, not just the credit decision. Four moments carry duties, and an agent touches all four.

  • What is asked. §1002.5 restricts inquiries into protected characteristics — race, colour, religion, national origin, sex, and marital status and childbearing intentions in defined circumstances — with narrow exceptions such as monitoring information on dwelling-secured loans. A free-form conversation is an inquiry engine.
  • Whether anyone was discouraged. §1002.4(b) prohibits statements to applicants or prospective applicants that would discourage a reasonable person, on a prohibited basis, from making or pursuing an application. The liability attaches to something said, with no application and no decision required.
  • When the application became complete. §1002.9 runs a thirty-day notification clock from the receipt of a completed application. An agent that gathers the last missing document at 2 a.m. has started that clock.
  • What the notice says. §1002.9(b)(2) requires the specific principal reasons for the adverse action — not a category, not a hint, and not a reference to a model.

Anchor on the regulation, not on the guidance. Supervisory guidance about AI and complex models in underwriting has been issued, withdrawn and reissued since 2023; §1002.9 has not moved in that time. In the EU, evaluating the creditworthiness of a natural person is Annex III high-risk under the AI Act — with the standalone high-risk obligations deferred to 2 December 2027 by the Digital Omnibus regulation adopted in mid-2026 — and Article 22 GDPR independently governs solely-automated decisions with legal or similarly significant effects. Design to the obligations; treat every date as movable, because they have all moved once.

STEP 2

A conversational agent is a regulated interviewer. Enumerate its questions.

The instinct is to instruct the model not to ask about protected characteristics. That is a prompt, and a prompt is a preference — see the instruction hierarchy for why it is the wrong layer for a control you must hold. Applicants also volunteer protected information unprompted, which no instruction prevents.

  • Work from a closed question set. Every question the agent may ask is an entry in a reviewed list with a stated purpose. The model chooses which to ask next and how to phrase it warmly; it does not compose new questions. Follow-ups that are not in the list go to a human.
  • Filter on the way in, not on the way out. Volunteered protected information gets stripped at intake, before it reaches the decisioning context at all. If it is in the context, you will be arguing about whether it was used; if it never entered, there is nothing to argue about.
  • Treat proxies as the same category. Neighbourhood, school, first language, name origin, a foreign address history. These are the variables an adverse-impact analysis will find, and an agent that reads unstructured documents ingests them whether or not you asked. Permission-aware, purpose-scoped retrieval matters here — see permission-aware retrieval.
  • Keep the required monitoring data on a separate path. Where collection is mandated, it goes to a compliance store, never into the context the decision is made from.
// The intake contract. The model picks from this; it does not extend it.
{ "question_id": "income.employment.gross_monthly",
  "purpose": "capacity",
  "required_for": ["installment", "revolving"],
  "phrasing": "model-composed, reviewed template",
  "storage": "decision_context" }

// Anything not in the registry:
{ "action": "route_to_human", "reason": "off-script inquiry" }
STEP 3

Discouragement is what helpfulness produces here.

An assistant trained to save the user time will tell a marginal applicant that they are unlikely to qualify and suggest they come back after six months of clean payments. That is a genuinely useful sentence in almost every other product, and in this one it is the discouragement risk in §1002.4(b) — said to a prospective applicant, before any application exists, with no adverse action notice to accompany it and no record that a decision was ever made.

  • Never let the agent predict the outcome. It may state criteria and it may explain the process. It may not say what will probably happen to this person, however kindly.
  • Make abandonment a measured, segmented outcome. Track where applicants stop, and compare drop-off across the geographies and demographics available to your fair-lending programme. A pattern where one group abandons at the same step is an adverse-impact signal arriving in a place no model test looks.
  • Let people apply anyway, always. The path to "submit this application and get a decision" must be reachable at every point in the conversation, including after the agent has explained that the criteria look tight.
  • Review the agent's marketing surface as advertising. §1002.4(b) covers advertising too, and a conversational pre-qualification widget is a statement to prospective applicants regardless of where your organisation files it internally.

This is the specific reason a human checkpoint placed only at the decision is misplaced in this domain: by the time the decision exists, three of the four regulated acts are already history.

STEP 4

Define "completed application" in code, because the agent decides when it happens.

The thirty-day clock starts on receipt of a completed application, and completeness is now determined by an agent's judgement about whether the pay stub it just parsed was sufficient. Left implicit, this becomes a timing violation that surfaces months later in an examination, from records that cannot reconstruct when completeness occurred.

  • Write the completeness predicate as a rule, not a prompt. A named list of required items, each satisfied by a verified artefact. The agent gathers; the rule decides; the transition is a logged event with a timestamp.
  • Emit the notice of incompleteness deliberately. Reg B gives a defined route for incomplete applications; an agent that simply keeps asking for documents forever has chosen neither route.
  • Make the clock visible in operations. Days-remaining on every open application, alerting before the boundary, because the failure mode is silent and dated.
  • Handle withdrawal honestly. An applicant who stops responding has not withdrawn; treat inferred withdrawal as a decision that needs a rule and a record, not a timeout.
STEP 5

The reason must name a factor the agent extracted — so measure extraction on denials.

Keep scoring in a declared model with enumerated variables, outside the agent. That much this playbook shares with underwriting agents, and it is the reason a black-box explanation problem never arises. The part specific to credit is subtler: the value of a declared variable now frequently comes from an agent parsing a document, and the notice must state the reason accurately, not merely plausibly.

  • Every extracted value is a named variable with provenance. Which document, which page, which model version, and the confidence. "Income insufficient" is a defensible reason only if you can show the income figure and where it came from.
  • Measure extraction accuracy on the denial population specifically. Aggregate accuracy is dominated by clean approvals. The cases that generate notices — thin files, irregular income, handwritten or foreign-language documents — are the hard tail, and the error rate there is the number that predicts wrong reasons and complaints. Sample and hand-score them monthly.
  • Route low confidence to a human before the decision, not after the notice. An abstention on extraction is cheap; a wrong principal reason on a mailed notice is not. This is calibration used as an operating threshold.
  • Keep the reason count small and specific. Disclosing a long list of marginal factors obscures the determinative one; the standard is the principal reasons, and a notice that reads like a model dump satisfies neither the applicant nor an examiner.
  • Never let the model write the notice text freely. Reasons come from a fixed catalogue bound to the declared variables. The agent may localise and format; generated prose in an adverse action notice is an unreviewed legal document.

Note the second notice regime sitting on top of this: where a consumer report contributed, FCRA adverse action duties apply as well, with their own content and their own timing. Two notices, one event, and an agent that touches the inputs to both.

STEP 6

Build the file the examiner will ask for, and a route to argue with it.

The request is always the same shape: reconstruct this application as of the day it was decided. Five artefacts answer it, and none of them can be assembled afterwards.

  • The full transcript, retained with the file, because three of the four regulated acts are utterances. This is the domain where the trajectory is the evidence, not a debugging luxury.
  • The variable values with provenance, and the model version that scored them — the same point-in-time reconstruction requirement as model risk management.
  • The completeness event and the notice date, so the clock is a record rather than an inference.
  • The fair-lending test results for the period, including drop-off by stage from STEP 3, not only outcome disparity on decided applications.
  • A working appeal path. An applicant who says the income figure is wrong should reach a human with authority to correct the variable and re-run, which is contestability with a defined remedy rather than a complaints inbox.

Before you build anything, take fifty real denial files and try to write the adverse action notice from the artefacts your proposed design would have produced. Most teams discover the determinative factor traces back to a document nobody stored, or a number nobody recorded a source for — and finding that on fifty files costs a week, while finding it in an examination costs a remediation programme. Ship the question registry, the completeness rule and the reason catalogue first; the conversational polish is the easy part and it is worth nothing without them. Related: collections and dunning agents for the far end of the same lifecycle, and KYC and AML onboarding for the identity checks that run alongside intake.