

A transaction monitoring model flags a payment. An investigator opens the alert, reads what the system produced, agrees with it, and closes the case.
Ask afterwards who made that decision and the answer is genuinely unclear. A human signed it, which satisfies the paperwork. Whether that human evaluated anything, or simply deferred to a score they had no way to interrogate, is a different question, and it is the one Sriramakrishna Vadlamudi thinks the industry keeps failing to ask.
Vadlamudi has spent close to a decade in enterprise financial services technology, moving from development work into anti-money-laundering compliance automation and the design of systems that deliberately keep people inside them. His objection starts with how the debate is usually staged.
"The debate over AI in financial crime compliance is often framed as automation versus human judgment, as if it's a dial to be turned one way or the other," he says. In his account the better systems do not turn the dial down as automation matures. They redesign where and how the human engages.
The distinction he draws is narrow and does most of the work in his argument. Automating data validation, case routing, even the drafting of narratives, is not the same as automating the decision. Those are different objects, and conflating them is how oversight quietly disappears while every procedure still appears to be followed.
He built a system on that principle. An automated case referral and routing tool he designed handles the data validation and routing that analysts previously did by hand, and by his account cut manual processing errors by roughly forty percent. The downstream effect was on time rather than accuracy alone, since fewer errors meant fewer cases needing rework, re-review or re-routing, which removed a recurring source of delay from the investigation queue.
What the tool deliberately does not do is dispose of cases. Structured checkpoints remain where a person has to review and approve before a case closes.
That was a design decision rather than a limitation, and he is clear about the difficulty of making it. There is no established playbook for where the line sits. Automate too little and the efficiency never arrives. Automate too much and human review degrades into a rubber stamp, which is worse than no review at all, because it produces a record of oversight that did not happen.
The failure mode he describes is specific and, on his account, common.
"The institutions getting this wrong tend to treat 'the model recommended it' as sufficient justification, which quietly erodes the very oversight regulators are relying on."
Which leads to the claim most likely to be argued with, and the most useful idea in his work.
Explainability, he says, is widely treated as a data-science requirement. He thinks that is a category error. It is a compliance control. "A model that can't produce a defensible explanation for an individual decision isn't just a technical shortcoming, it's an audit gap."
Reframed that way, the requirement changes shape. A data-science standard for explainability is satisfied when the technique is sound. A compliance standard is satisfied only when a specific person, reviewing a specific alert, can act on the explanation. Those are not the same bar, and the second is considerably harder.
It is also the problem Vadlamudi worked on directly. Explainability techniques exist in abundance in the data-science literature, but rendering them usable by a compliance investigator rather than a data scientist is a separate discipline. He built a feature-attribution approach that surfaces which specific factors drove a model's score, in a form a reviewer can actually evaluate, so that the human in the loop has something to exercise judgment on rather than a number to accept or reject blindly.
Without that, he argues, human oversight is structurally unable to function, whatever the procedure says. A reviewer confronted with an opaque score can defer or refuse, and neither is judgment.
The shift he considers most consequential is not better detection at all. It is agentic AI beginning to take on multi-step investigative and drafting work inside compliance operations, where a system gathers evidence, assembles a narrative and recommends a disposition rather than producing a single flag.
Existing oversight models do not transfer to that cleanly, because they were built around human investigators. His response has been to argue that each intermediate step an agent takes needs the same auditability as an analyst's actions would, rather than only the final output being examined. An agent that reaches a defensible conclusion through steps nobody can reconstruct has not been overseen. It has been trusted.
He has also looked at the question from the systems side, modelling how case backlog, cycle time and the risk of breaching service commitments behave as volume grows. The point of that work is to distinguish where automation genuinely relieves pressure from where human capacity remains the binding constraint by necessity rather than by policy. It is a useful corrective to the assumption that every queue problem is an automation problem.
His recommendation to institutions building in this area is about sequence, and it is unglamorous.
Human checkpoints should be first-class design requirements from the beginning, not compliance overlays fitted once a model already works. Retrofitting them afterwards is more expensive, and considerably less persuasive to an examiner asking how a decision was actually reached.
The institutions that manage this, in his view, will not be the ones with the most capable models. They will be the ones that can still answer the question the whole apparatus exists to answer: who decided, on what basis, and can they show it.