AI Governance Risk Benchmarking: Fragmented Ownership Despite Established Committees
- Most banks have established AI governance committees to manage the integration of generative AI (GenAI) and large language models (LLMs), though ownership of these frameworks remains fragmented across...
- The report indicates that financial institutions are building governance structures for GenAI in parallel with existing model risk management processes.
- This fragmentation affects how banks handle model validation and the mitigation of operational and reputational risks associated with AI deployment.
Most banks have established AI governance committees to manage the integration of generative AI (GenAI) and large language models (LLMs), though ownership of these frameworks remains fragmented across different departments, according to the Model Risk Benchmarking 2026 report.
The report indicates that financial institutions are building governance structures for GenAI in parallel with existing model risk management processes. While the creation of dedicated committees is widespread, the lack of centralized ownership creates a gap between high-level oversight and the technical execution of risk controls.
This fragmentation affects how banks handle model validation and the mitigation of operational and reputational risks associated with AI deployment.
Fragmented Ownership of AI Governance
The Model Risk Benchmarking 2026 data shows a disconnect between the existence of governance committees and the actual ownership of AI risk. Many banks have formed committees to provide a veneer of oversight, but the responsibility for managing those risks is split between the Chief Information Officer (CIO), risk management teams, and individual business units.
This split ownership complicates the implementation of consistent standards for model risk. When governance is fragmented, the process of identifying and escalating risks becomes slower, potentially leaving gaps in how GenAI tools are monitored in production environments.
Integration of GenAI and Model Risk Management
Banks are currently treating GenAI governance as a separate but parallel track to traditional model risk management. Traditional model risk focuses on quantitative models used for credit scoring or capital requirements, whereas GenAI introduces new variables such as “hallucinations” and non-deterministic outputs.
According to the benchmarking report, the primary challenge lies in adapting existing model validation techniques to fit the nature of LLMs. Traditional backtesting—the process of applying a model to historical data to see how it would have performed—is difficult to apply to generative AI, which produces unique responses rather than a single numerical prediction.
To address this, institutions are attempting to merge GenAI oversight into their broader operational risk frameworks. This includes monitoring for reputational risk, where an AI-generated response could lead to public relations failures or regulatory scrutiny.
Technical Challenges in Model Validation
The shift toward LLMs has forced banks to rethink the role of the model validator. The report highlights several specific technical hurdles:
- Backtesting Limitations: The lack of static outputs in GenAI makes standard backtesting insufficient for ensuring model stability.
- Model Drift: LLMs can evolve or degrade in performance over time, requiring continuous monitoring rather than a one-time validation at the point of deployment.
- Data Privacy: Ensuring that sensitive bank data does not leak into the training sets of public LLMs remains a primary operational risk.
Because these challenges differ from those found in traditional financial modelling, the “parallel” approach to governance has become the default strategy for most institutions.
Impact on Operational and Reputational Risk
The Model Risk Benchmarking 2026 findings suggest that the stakes for GenAI governance extend beyond technical accuracy. Reputational risk is cited as a critical concern, as the public-facing nature of many AI tools means that a single incorrect or biased output can result in immediate brand damage.
Operational risk is also heightened by the “black box” nature of many LLMs. When a model makes a decision or provides a piece of information, the inability to trace the exact logic path—known as explainability—creates a conflict with regulatory requirements that demand transparency in financial decision-making.
Banks are currently attempting to solve this by implementing “human-in-the-loop” systems, where a human reviewer validates the AI output before it reaches a client or affects a financial transaction.
