Policy · August 16, 2026 · 6 min read
What We Told RBI About Data Governance for AI
RBI's draft data governance guidance and its draft model risk guidance govern two halves of the same pipeline. Neither yet reaches the seam where data becomes an automated decision. We filed our comments; here are the seven things we asked for.
On July 15 the Reserve Bank published its draft Guidance on Regulatory Expectations for Data Governance and gave the industry until August 17 to respond. It arrived three weeks after the draft Model Risk Management guidance, applies to the same eleven classes of regulated entity, and comes from the same Operational Risk Group in the Department of Regulation. Read together, the two drafts govern the two halves of every automated decision a bank makes: the data on one side, the model on the other.
We filed our comments on August 16. This post is the public version: seven specific asks, and the one conviction underneath them.
The theme: governance has to reach the decision
The draft is good. It asks for a lifecycle view of data from origination to disposal, a single source of truth, documented metadata and lineage, named data owners, stewards and custodians, quality metrics that go to a board committee, and controls on data shared with third parties. It aligns itself explicitly with the DPDP Act and Rules. Para 1 names “automated decision-making processes” as one of the reasons the guidance is needed.
What it does not yet do is follow the data to where automated decisions actually happen. Lineage in the draft runs “from origin through aggregation, transformation, and usage to its final destination.” For an AI system the final destination is not a warehouse table. It is the moment an agent retrieves three records, assembles a context, sends it to a model, and acts. Every property the draft cares about (lineage, quality, purpose limitation, third-party control) is most exposed at exactly that moment, and it is the moment the draft’s language stops just short of.
Our comments converge on one adjustment: data governance has to reach the point of automated use, the decision, not stop at the store. Everything else follows.
What we asked for
1. Extend lineage to the automated decision (paras 5(8), 33(ii), 43, 47-51). The draft requires metadata to flow with data to downstream systems and to be updated on transformation. We asked that “downstream usage” explicitly include the specific automated decision the data produced: the records retrieved, the context assembled, the tool outputs consumed. Captured contemporaneously, because provider-side model updates destroy any hope of reconstructing it later, and paired with configuration provenance (model version, prompt version, tool manifest) so a decision is reconstructable from both what it read and what read it.
2. Third-party safeguards must cover inference-time egress to model providers (paras 60-63, 15(4), 39). Chapter VI is about data shared with third parties. In an agentic system the largest and least-governed sharing of customer data is the prompt: the context and tool inputs sent to a hosted model to get an answer, continuously, thousands of times a day, and often offshore. We asked that the guidance say so, and require three things for those flows: tokenise or pseudonymise identifiers before the external call (para 39 already lists tokenisation as an expected control), verify where any retained raw identifier actually lives without exposing it, and keep a per-decision record of which data category went to which provider. This is the point in our letter we care most about.
3. Tie data quality to model monitoring and India-calibrated fairness (paras 56-59). Para 57 says quality gaps must not compromise decision-making. For AI, quality defects do not look like errors; they look like slow drift and biased outputs. We asked that quality metrics be visible to the model-monitoring function and escalatable as model-risk events, and that fairness be assessed on production decision distributions across India-relevant dimensions and their proxies.
4. Purpose limitation, evidenced at the point of access (paras 25(ii), 28(i), 34-35, 47-48, 52(i)). The draft records permitted usage in metadata and asks custodians to enforce access in line with it. It does not yet ask for evidence of the actual access act. An agent that reads data for a purpose it was never collected for is where most purpose drift happens in automated systems. We asked that REs retain evidence of which system or agent accessed which data, for what purpose.
5. A named human owner for every automated data consumer (paras 30-31, 33(iv), 47). The role architecture presumes human actors. An agent is not a data steward and cannot bear accountability. We asked that the required role mapping name an accountable human for each automated or agentic consumer of data, and that such consumers be inventoried as data-accessing entities.
6. Make lineage evidence tamper-evident and independently verifiable (paras 15(1), 20, 44(3)). Integrity, auditability and traceability are already principles in the draft. We asked for a stated preference that lineage and quality evidence be append-only or cryptographically verifiable, so a supervisor can check the lineage of a decision without trusting the RE’s assertion that the record is complete.
7. Read the two drafts together, and give both a graded timeline (paras 3, 4, 6(i)). Neither draft addresses the seam between them; neither has an effective date. We asked for an explicit read-across and a phase-in that is proportionate for smaller REs and consistent across the two guidances, so institutions implement once rather than twice.
Why we’re saying this in public
Because the two drafts are the shape of what Indian financial regulation is going to ask of AI systems for the next several years, and the seam between them is where the interesting failures will happen. A model that is perfectly validated on data whose lineage stops at the warehouse, or a data programme that is perfectly governed up to the moment an agent reads it and calls an API, is the kind of gap that looks fine in a policy review and terrible in an inspection.
We build for this. Sakshi records what an agent read, from which system, for what purpose, and what it decided, with identifiers tokenised before anything leaves the institution’s environment; it verifies where retained identifiers live without seeing them; and it produces the evidence pack that maps to the clause. Not because we predicted the draft, but because we think both drafts are right about what governing AI in finance requires.
The full text of the draft guidance is on rbi.org.in under Connect 2 Regulate. Our comment on the draft Model Risk Management guidance is here. If you are an RE working out what the two drafts mean together, we should talk.