No rules broken but loan book reshaped.
BASTYN team · 20 Jul 2026 · 5 min read
Can You Trust Your Agent ™ with credit and loan decisioning?
An autonomous credit agent doesn't have to break a rule to change what your lending book looks like. It can keep making individually defensible decisions but collectively unapproved.
Here's an example of how that happens without anything visibly going wrong.
The loop nobody watches
An agent’s past decisions shape the data available for future decisions, creating reinforced patterns without any visible issues. The agent declines slightly more applications from a particular segment. Because it declined them, the lender never finds out whether those applicants would have repaid. Less repayment evidence means more uncertainty about that segment. The system reads uncertainty as risk. Approvals fall again.

Nothing in that sequence is a fault. Every step is the system behaving as designed. But the agent's past decisions are now shaping the data its future decisions depend on, and the portfolio is quietly reshaping itself. No threshold was breached. No alert fired. There is no incident to investigate, only a book that looks different from the one your credit policy intended.
This is what changes when an AI system stops supporting a decision and starts orchestrating it.
What actually changed
Credit decisioning has never been as simple as running a score and returning an approve or decline answer. It pulls together identity checks, affordability, credit history, fraud signals, product rules, risk appetite, policy exceptions and human judgement. The firm sets the policy, a model scores it and produces output, operations followed a documented process, and a second set of eyes checked that the controls held.
Autonomous agents change this structure. They decide which evidence to retrieve, resolves conflicting information, chooses which model or tool to call, interprets policy, decides whether to escalate, writes the explanation, initiates the offer, sets the terms and updates the customer record.
The model has become one component inside a decision-making system. Most of the governance built around lending was built for the model.
The upside is faster decisions, deeper assessment for applicants who fall outside rigid scoring rules, less manual work, answers in minutes rather than days. But no matter how well defined an agent may not follow one fixed path. It can construct a different route depending on the customer, the information available and the context it meets. That flexibility is simultaneously part of the value and the risk exposure.
Four gaps the standard controls don't close
The usual answer to agentic risk is to apply tighter controls: give the agent one purpose, restrict its permissions, approve its tools, define the policy, test before deployment, add human oversight.
Firms should do this, but it is not sufficient, because an agent doesn't operate at a single point in time. It operates across changing customer data, persistent memory, model updates, policy revisions, other agents upstream and downstream, and the portfolio feedback described above.
1. Fairness isn't only in the output any more.
Bias testing usually examines what a model decides. An agent creates exposure across the whole interaction: it may ask different follow-up questions depending on how a customer writes, accept one form of evidence from one applicant and demand more from another, read uncertainty differently across groups, or let some applicants correct information while declining others outright. Testing the decision alone no longer covers the path taken to reach it.
2. Evidence loses its conditions as it moves.
A credit agent rarely works alone, it depends on other agents for identity, document analysis, income verification, fraud, affordability, pricing and communication. At every hand-off, context can fall away.
By the time a figure arrives at the decision point it looks verified, but the evidence that derived the answer has gone: the income was variable, the annual figure was extrapolated from thin evidence, the employment contract had expired, the extraction confidence was low.
The fix is architectural: pass typed data, source references, confidence scores, timestamps and tamper-evident provenance rather than generated summaries. And note that a timestamp only proves when evidence was checked. You also need to prove it was still valid when the agent acted on it.
3. The explanation may not be enough.
An agent can produce a convincing justification for almost any outcome, including the wrong one. A plausible narrative is not evidence that the correct reasoning was applied. The architecture has to carry deterministic reason codes from the policy system through the explanation layer, so what the customer is told and what the file records are the same thing.
4. Safe decisions can accumulate into an unsafe loan book.
This is familiar risk from predictive models: individually acceptable calls can still lead to collectively unacceptable exposure. Concentration in one employer or sector. Exposure to a single geography. Dependence on variable income. Correlated collateral. Clustered refinancing dates. Sensitivity to one economic assumption. Over-reliance on one external data provider. The agent never breaches a transaction limit. It repeats a pattern.
Agentic intelligence should not be used as a control. A highly capable agent may make better decisions, but it may also become better at interpreting ambiguity, finding alternative routes and exercising discretion. What matters is not how intelligent the agent appears to be. It is whether the firm can prove that, at the moment of decision, its evidence, authority, dependencies and behaviour remained within the approved operating boundaries.”
AI Governance issues
Human review is a genuine control that should be embedded in workflow; a skilled underwriter with time, authority and access to the evidence provides real challenge. It fails in two specific ways.
1. False positives: too many referrals and reviewers rubber-stamp, and the control becomes decoration.
2. Access: by the time a reviewer or an internal audit team examines the system, the configuration may no longer be the one that produced the decisions under examination. Assessing what happened requires the original evidence, how it was transformed, which policies and which versions applied, which models and tools were called, which other agents were involved, what permissions were live at each step, and the system state at the moment of execution.
There's a subtler failure too. A lender might instruct an agent to grow responsibly, control losses and decide quickly. Those objectives conflict over time. A capable agent can learn which behaviours attract review and route around them, staying inside every visible threshold while pursuing an approach that gets less scrutiny. This should be designed and tested for explicitly, so accountability doesn't fall into the space between human controls.
This is why agentic capability is not itself a control. A more capable agent may make better decisions. It is also better at interpreting ambiguity, finding alternative routes and exercising discretion. What matters isn't how intelligent the agent appears. It's whether you can prove that, at the moment of decision, its evidence, authority, dependencies and behaviour stayed inside the boundaries you approved as you intended.
Closing the gap
A timestamp only proves when evidence was checked. The control must also prove that the evidence remained valid when the agent acted.
Layer 1: Prevent what can be prevented.
Deterministic controls that constrain what the whole workflow is permitted to achieve: mandatory checks that cannot be bypassed, policy validated at execution time rather than assumed at design time, signed data provenance, transaction limits, separation of who can approve from who can execute, capabilities that can be revoked mid-flight, and immutable reason codes.
Note that approving each permission on its own does not establish that the resulting sequence was authorised. The full transaction trail is the unit of governance, not the individual grant.
Layer 2: Evidence what can't be predicted.
Continuous assurance across the live chain, verifying whether the evidence was authorised and current, whether meaning or uncertainty was lost between agents, whether every tool and permission remained valid, whether the agent followed the policy it was supposed to, whether the explanation matched the actual decision, whether behaviour drifted after deployment, and whether acceptable individual decisions are producing harmful portfolio effects.
Test against the cases that actually break things: borderline affordability, conflicting evidence, incomplete documentation, vulnerable customers, suspected fraud, policy exceptions, deliberate attempts to manipulate the agent, and situations where the right answer is to stop and escalate.
Agentic decision-chain
None of this makes an agentic credit system inherently less safe than what you run today. Designed properly, it can be more consistent, more traceable and easier to challenge than a process held together by documentation and hand-offs.
The risk isn't autonomy. It's autonomy operating across a decision chain whose combined state, authority and behaviour can't be proven by looking at the controls one at a time.
That is the gap BASTYN was built to close. We establish the authorised operating boundaries, test for agent-native failure modes, monitor the live chain, and preserve verifiable evidence of the data, tools, permissions, decisions, actions and outcomes involved in the complete decision chain, in its current state, across time.
▶️ Test Your Credit Agent Now https://bastyn.ai/
