At a glance
- Independent approval does not prevent several banks from sharing the same hidden weakness.
- When circumstances change, human sign-off can become an escalation bottleneck.
- Governance needs signals that conditions have changed, authority to suspend a skill and visibility of shared dependencies.
It is challenging to think through AI risk and governance. At Beacon58, we use thought experiments to bring them to life.
Here's one.
Skills are becoming big business, with downloads running into the millions. A skill is a packaged set of instructions that tells an AI agent how to do a specific job. Businesses are adopting them centrally, delighted at the efficiency and productivity they bring.
Imagine several Gulf banks using the same financial-crime skill. Their existing systems still generate alerts. The skill takes over the investigation that follows: gathering a customer's history, weighing possible explanations, writing the case file and recommending an outcome.
A human signs off every case. Each bank has tested and approved the skill independently.
The skill is designed to question sudden changes in payment routes, unfamiliar counterparties and activity that does not fit a customer's history. These are sensible reasons to investigate. The banks are thrilled as cost-to-income ratios are falling and shareholders are pleased.
Then a regional conflict changes what normal looks like. Shipping firms are forced to use unfamiliar ports, businesses replace suppliers they can no longer reach, and families send money through different corridors. Much of this is legitimate, but sanctions evasion is rising too, so banks cannot simply dismiss the alerts.
Case volumes multiply across the banks in the same week. Investigators receive files recommending escalation faster than they can properly assess them. Overturning a recommendation takes evidence and time, and clearing a case that later proves criminal is a personal risk.
Confirming the need for escalation is low risk, so sign-off increasingly means agreeing with the skill. The banks cannot hire more investigators, because they are all short of the same people, and there is no time to test a recalibrated skill.
The consequences spread beyond individual cases. A restriction on one exchange house can disrupt remittances for thousands of workers. A customer who tries another bank meets the same skill, which reaches the same conclusion. Some people move to cash or informal channels, which makes them look riskier at the next review.
On the banks' dashboards, performance appears to improve: more alerts escalated, faster case preparation and more consistent decisions.
No agent has gone rogue, and no skill has malfunctioned. A skill created and tested in stable conditions is being applied to a world that has changed.
If you have approved an AI skill, consider these questions: How would you know that the conditions it was calibrated for no longer hold? Who could suspend it within the week? How many other institutions in your market are running the same one?
What do you think, fantasy or plausible? How are you applying AI governance to address future scenarios? If you need help, please get in touch.