Policy · July 18, 2026 · 5 min read
What We Told RBI About Governing AI Agents
RBI's draft Model Risk Management guidance is the most consequential AI regulation Indian finance has seen. We filed our formal comments this month. Here are the seven things we asked for, and the one theme underneath all of them.
On June 24, the Reserve Bank published its draft Guidance on Regulatory Principles for Model Risk Management, 2026 and gave the industry a month to respond. It is, quietly, the most consequential AI regulation Indian finance has seen. It applies to every commercial bank, small finance bank, UCB, NBFC across all layers, ARC, and credit information company in the country, and its definition of “model” (algorithms, analytics, interfaces, applications, decision-based rules, irrespective of whether such tools are recognised as models by the RE) covers essentially every AI system a financial institution runs, including the agentic ones being deployed right now.
We filed our formal comments this month. This post is the public version: seven specific asks, and the one conviction underneath all of them.
The theme: the unit of governance is shifting
Model risk management grew up in a world where a model was a scoring function: inputs in, number out, validate annually. That world is ending. Models are becoming components of autonomous systems, agents, that retrieve context, invoke tools, chain reasoning, and act. Two banks can deploy the same model and carry radically different risk, because risk now lives in what the system around the model is allowed to do.
Our comments converge on one adjustment: as models become agents, the unit of governance shifts from the model to the decision. Everything else follows from that.
What we asked for
1. Inventory the deployed system, not just the model (paras 21-22). The draft’s inventory mandate is exactly right. “No model is used, relied upon, or deployed unless it is part of inventory” is the strongest sentence in the document. We asked that the minimum attributes include autonomy level, operational footprint (what systems and money the deployment can touch), and configuration provenance: the model version, prompt version, and tool manifest actually in effect. Without those, two identical-looking inventories can hide a tenfold difference in exposure.
2. Traceability requires decision-time capture (para 57). The draft asks for documentation enabling “traceability, reproducibility, and auditability” of AI models. We made a technical point that deserves to be a regulatory one: for AI systems, post-hoc reconstruction is a fiction. Provider-side model updates and shifting knowledge bases destroy reproducibility within weeks. The only artifact that survives is a contemporaneous, tamper-evident record of the decision (context, intermediate steps, tool outputs, model version) captured at the moment it happened. This is implementable with current technology; we know because we implement it.
3. Make human oversight measurable (paras 60-63). The draft is ahead of most jurisdictions in naming automation bias and decision fatigue. We asked for one strengthening. Require institutions to measure the loop: override rates, review time-on-task, escalation outcomes. A near-zero override rate on a high-volume approval queue is not evidence of a good model; it is evidence of a reviewer who has stopped reviewing. Risk-tiered deep review of fewer decisions protects customers better than shallow ratification of all of them.
4. Drill the kill switch (paras 60(ii), 43). The draft requires kill-switch arrangements. We asked that they be periodically drilled, with attestation records (propagation time, scope verification, fallback behavior), exactly as BCP testing regimes already demand. An untested kill switch is a compliance claim, not a control.
5. Record the provider’s model version with every decision (paras 45-56). For foundation models consumed via API, the institution controls neither the weights nor the update cadence. The model that made a disputed decision may not exist by the time of the dispute. Contractual accountability cannot substitute for decision-time evidence.
6. Give the guidance a timeline (para 64). The draft specifies no effective date or phase-in. A framework this comprehensive, reaching down to base-layer NBFCs and rural co-operative banks, is more credible with a staged runway than with silent immediacy. We proposed roughly six months for framework and inventory, twelve for the engineering-dependent requirements on customer-affecting AI systems, with proportionate expectations for smaller institutions.
7. Calibrate fairness testing to India (para 54(3)). Bias assessment that checks Western benchmark dimensions will pass models that fail Indian customers. Fairness testing here has to cover India-relevant dimensions and their proxies in credit data (name, locality, occupation patterns), and it has to run on production decision distributions, not just pre-deployment benchmark sets.
Why we’re saying this in public
Because the direction of travel is set. The comment window closes July 24; final directions are expected within months. When they land, procurement begins, and institutions that treated this draft as a reading exercise will discover that “traceability, reproducibility, and auditability” is an engineering program, not a policy PDF.
We build for this. Every ask in our letter is something Sakshi produces today: the agent inventory with autonomy tiers and provenance, the tamper-evident decision records, the oversight telemetry that distinguishes review from ratification, the drilled kill switch with signed attestations, the clause-mapped evidence report. Not because we predicted the draft, but because we think the draft is right about what governing AI in finance requires, and someone has to make it real.
The full text of the draft guidance is on rbi.org.in (Press Release 2026-2027/528). If you’re an RE working out what it means for your AI programme, we should talk.