Artificial Intelligence

‘AI Governance Must Be Engineered Into Banking Systems’: Maveric Systems’ Kishan Sundar

AI Governance Must Be Engineered Into Banking Systems: Maveric Systems’ Kishan Sundar

Written By : Analytics Insight

Authored by Kishan Sundar

Every conversation about AI governance in banking eventually arrives at the same place: regulation. Boards ask what the regulator will require. Compliance teams map frameworks to model inventories. Risk committees debate disclosure obligations. While all this matters, what often gets overlooked is where governance is determined. Governance is built into the system itself, not documented in a policy.

By the time a regulator asks a bank to explain a model's decision, that explanation either exists or it doesn't. It cannot be manufactured retroactively with a good compliance narrative. If the lineage wasn't captured when the decision was made, if the data that fed the model wasn't tracked, if the reasoning wasn't logged, there is nothing to audit. 

This is the uncomfortable truth many financial institutions are learning as their AI programs mature: governance is not a layer added once a model works. It is a property of how the model, the data, and the surrounding systems are engineered from the outset.

Trust is Engineered, Not Audited

For most of the last two years, generative and agentic AI in banking was treated as a technology experiment. Success was measured by whether the model worked in a controlled pilot, not whether it could operate reliably within the complexity of a production banking environment.

As a result, many AI initiatives stalled after the pilot stage. The challenge was rarely model accuracy. Instead, pilots were built on clean data, focused on narrow use cases, with little integration with core banking systems and no clear approach to explaining decisions to regulators, auditors, or customers.

Today, that approach is no longer sufficient. AI has evolved from an experiment into a business capability. It underwrites credit decisions, screens transactions for fraud, drafts customer communications, and acts on behalf of both the institution and the customer. At this stage, success depends less on the model itself and more on the system around it.

That means treating explainability, traceability, and oversight as engineering requirements rather than compliance activities. Before a single line of model code is written, institutions should ask: How will this decision be explained? Who validates it before and after deployment? What happens when it is wrong? Where does the evidence live?

The Contextual Data Catalog as Foundation

If governance is engineered, it starts with data specifically, with the difference between data that is merely available and data that is decision-ready.

Most banks have no shortage of data. What they lack is a contextual data catalog: a structured, living map of not just where data sits, but what it means, where it came from, how it's governed, and which rules, SOPs, and regulatory obligations apply to it. This is less a data dictionary than a knowledge fabric, one that stitches raw data together with the processes, policies, and operational knowledge that give it meaning inside a specific banking workflow.

This distinction matters enormously for explainability. A model that can point to the fields it used is not the same as a model that can explain why those fields were relevant and what lineage they carry from source system to decision. Explainability through lineage means every decision can be traced backward, not just to a dataset, but to the specific record, transformation, policy, and approval chain that produced it. Without that lineage, "explainable AI" is a marketing phrase. Explainability becomes a natural consequence of decision lineage rather than an exercise performed after deployment.

This is also where a subtle but important shift in thinking is happening across the industry: the priority is no longer finding the single "best" model. Context is more valuable than choosing the best model - a well-governed institution can swap models as better ones emerge. Still, it cannot swap out years of accumulated, structured context about its own data, processes, and regulatory environment. Enterprises should govern their knowledge and context as a durable asset, not chase whichever model tops this quarter's benchmark while treating outputs as the only thing worth governing.

Safe AI by Design

Once the contextual foundation is in place, the platform itself must enforce security, fairness, explainability, human oversight, auditability, and compliance. In a banking context, this translates into concrete engineering commitments woven directly into the platform. Security is built around data access controls, encryption, and model isolation, so sensitive financial data never leaks across use cases or tenants. Fairness is tested systematically across protected attributes and customer segments, rather than assumed to be balanced because the training data "looked balanced." Explainability is designed into the model interface itself, so every output carries its reasoning rather than requiring a separate investigation afterward. Human-in-the-loop checkpoints are placed deliberately at the decisions that matter most- credit denials, large transactions, account actions – rather than applied uniformly everywhere or nowhere. Audit trails are generated automatically as a side effect of the system running, not compiled manually after the fact. And compliance is built in by default, with regulatory constraints encoded as guardrails the system cannot bypass, rather than rules a user is trained to remember.

Model Validation and Monitoring

Financial institutions have long relied on model risk management for statistical and credit models. While those principles still apply, generative and agentic AI demand additional safeguards because their behavior is less deterministic and often relies on external knowledge sources.

That means validating AI from multiple angles. White-box testing examines a model's internal logic where possible, while black-box testing evaluates its behavior through inputs and outputs—essential for foundation models whose internals are opaque. For RAG systems, testing must also verify that responses are grounded in retrieved evidence rather than fabricated by the model.

Validation does not end at deployment. Models must be continuously monitored for bias and drift, as changes in customer behavior, market conditions, or fraud patterns can cause a once-reliable model to become inaccurate or unfair over time.

Agentic AI Increases Governance Complexity

If governing a single model is demanding, governing agentic AI is an order of magnitude harder, and it's where most institutions' current governance frameworks start to show their age.

Agentic architectures involve multiple specialized agents: one retrieving data, another applying policy logic, another drafting a response, and another executing a transaction, coordinated through an orchestration layer to complete a task no single model could handle alone. This introduces governance questions that don't exist in single-model systems: How is a decision attributed when it emerges from five agents interacting rather than one model responding? How are conflicts between agents resolved when they disagree, and is that resolution itself logged and explainable? At what point does the orchestration layer escalate to a human, and how confident must the system be before acting autonomously versus requesting approval?

Traceability is the foundation of trust in these systems, and it becomes structurally harder to maintain as more agents, more handoffs, and more autonomous decisions enter the chain. Multi-agent systems need orchestration layers that not only route tasks efficiently but also log every handoff, every piece of context passed between agents, and every point where one agent's output changes another's behavior. Without that, a bank cannot answer a basic question a regulator will eventually ask: which system, exactly, made this decision, and why?

This is also where human escalation paths need the most deliberate engineering. In a single-model system, human review can be inserted at one obvious checkpoint. In a multi-agent system, the institution must decide, agent by agent, where autonomy ends and human judgment begins, and build that logic into the orchestration layer rather than leaving it to case-by-case discretion.

Engineering the Future of AI Governance

The next phase of AI adoption in banking will be defined by an institution's ability to operationalize trust at scale. Delivering contextual data catalogs, safe-by-design architectures, rigorous validation, and traceable agent orchestration requires platform engineering.  As AI becomes embedded across critical banking functions, governance must be part of the platform itself.  For most banks, the challenge is ensuring AI can operate safely within decades-old infrastructure.

This is precisely why banking-focused technology partners are playing a central role in how global institutions operationalize responsible AI. They provide the engineering foundations that enable institutions to integrate AI with legacy core banking systems, enforce governance consistently, and scale responsibly across the enterprise.

In the coming years, competitive advantage in banking will be determined less by who adopts the latest model first and more by who can deploy AI responsibly across thousands of decisions every day. Institutions that treat governance as an engineering capability will help define the standard that the rest of the industry follows.

How Geopolitical Sanctions are Changing the Role of Crypto in Cross-Border Payments

Crypto Prices Today: Bitcoin Reclaims Above $80,000, Largest Weekly Surge; Solana Leads Weekly Gains Past 35%

Arthur Hayes Buys $1.17M ETHFI After Selling Token at a Loss

How Stablecoins are Changing the Future of Digital Payments, Global Money Transfers

BitMart Reconsiders Full Shutdown as Restructuring Plan Takes Shape