Out of Scope
Sixty-eight days apart, two financial regulators looked at generative AI. One put it outside the rules. The other wrote it in. I've been trying to work out what that means for anyone who has to build.
Provenance
Two documents landed this year, sixty-eight days apart.
One is American. It takes the most consequential technology in modern banking, looks straight at it, and sets it aside — not out of negligence, but out of a sort of honesty I’ve come to find more unsettling than negligence would have been.
The other is Indian. It reaches the opposite conclusion about the same technology, in the same quarter, and barely anyone has noticed the two are in conversation.
I can’t make them agree. What follows is an attempt to work out why — and what it means for the people who have to build inside the answer.
I should say up front that I’m not one of those people. I only think I might be standing next to them.
⚠️ Disclaimer: A note on where I’m standing: I don’t work in a bank. What follows is assembled from primary documents — supervisory guidance, regulator drafts, industry research — plus one relevant scar. I’ve sourced every figure. The conclusions are my current best guess and I’d revise them tomorrow given a good reason. If you sit inside model risk or a bank platform team, the last section is for you.
I. The question
I’ve been stuck on a question for a few months, and it started idle.
What’s the last thing AI touches?
Not “what will AI disrupt” — everyone has written that essay, including me. The inverse. Run the tape forward, watch the water rise: what’s the last rock still above the surface?
My first answer was the obvious one. The technically hardest domains — the arcane, the specialised, the twenty-year expertise moats.
That thesis hasn’t survived contact with the last three years. Complexity is exactly what these models are good at. They read the contract, write the code, pass the exam. If difficulty were the wall, the wall would be down.
So I flipped it:
AI arrives last wherever being wrong is expensive.
Not hard. Expensive. Where a mistake has a consequence, the consequence has an owner, and the owner has a name and a job and a mortgage.
II. The evidence that it isn’t the model
MIT’s NANDA initiative studied 300 public AI deployments, surveyed 350 people and ran 150 interviews: roughly 95% of generative AI pilots produce no measurable P&L impact. Gartner’s numbers rhyme — 89% of AI agent pilots never reach production, and they forecast over 40% of agentic AI projects cancelled by end-2027. S&P Global found the average organisation scrapped 46% of its AI proofs-of-concept.
If the models were the constraint, the failures would cluster around capability. They don’t. Here’s MIT’s breakdown of where successful AI effort actually goes:
10% algorithms. 20% technology and data infrastructure. 70% people and process.
Ninety percent of the work is not the model. And the ~11% of agent pilots that do reach production return, on Gartner’s estimate, 171% ROI — so this isn’t a story about the technology being oversold. It’s a story about the bottleneck sitting somewhere other than where everyone is digging.
The pilots aren’t dying in the lab. They’re dying in a room, in a meeting, in a conversation nobody was equipped to have.
III. So I went looking for the people who’d solved it
If the frontier is about governing consequence, go study whoever is best in the world at governing consequence. Which is banking.
American banks have been formally, supervised-ly managing model risk since 2011 — SR 11-7. Fifteen years of accumulated discipline. An entire profession of model validators.
And on 17 April 2026, the Fed, the OCC and the FDIC replaced it with SR 26-2.
I sat down expecting a masterclass.
IV. Footnote 3
“Generative AI and agentic AI models are novel and rapidly evolving. As such, they are not within the scope of this guidance. Nonetheless, a banking organization’s risk management and governance practices should guide the determination of appropriate governance and controls for any tools, processes, or systems not covered in this document. However, the principles described in this guidance apply to traditional statistical and quantitative models and non-generative, non-agentic AI models.”
Traditional statistical models: in scope. Non-generative, non-agentic AI: in scope. The exact category of system every large bank is racing to deploy: explicitly carved out.
To be fair to them, the reason given is the honest one. Novel and rapidly evolving. The agencies are declining to pretend they have a settled framework for something that changes shape quarterly. An RFI is promised. The guidance is non-binding anyway — no enforceable standard, aimed principally at institutions above $30 billion in assets.
So my reaction isn’t how could they. It’s closer to: the honest answer and the useful answer may not be the same answer here, and I don’t know what you do about that.
V. I am not the first person to notice this
Chris Stanley, banking industry practice lead at Moody’s, wrote about SR 26-2 in American Banker in June, and his framing is better than mine:
“SR 11-7, the replaced guidance, was designed to fight the risk of using bad models. The new guidance, SR 26-2, recognizes the risk of not using models.”
That’s the sentence I wish I’d written. The old regime’s failure mode was a bad model shipping. The new regime’s failure mode is a good model not shipping while your adversaries ship theirs.
He also says the quiet part, on behalf of practitioners I don’t have access to:
“Traditional pre-deployment validation is not just harder to apply to these tools; it’s nearly impossible. You cannot fully validate a model whose behavior evolves after deployment, whose inputs shift with the threat environment and whose applications are determined by the ingenuity of the people using it.”
His conclusion: “The footnote is not a reprieve. It is a challenge to adapt faster than the rules can be written.”
I think he’s right. And eight days after he published, the story got stranger.
VI. Meanwhile, 8,000 miles away
On 24 June 2026, the Reserve Bank of India released a draft Guidance on Regulatory Principles for Model Risk Management, open for public comment until 24 July.
It does the opposite of what the Fed did. AI models are squarely in scope — with provisions written specifically for their behavioural characteristics — and the applicability list is expansive: “Commercial Banks (including Foreign Banks)”, payments banks, co-operative banks, NBFCs across all four layers, the all-India financial institutions.
So:
The Fed, OCC and FDIC looked at generative and agentic AI and put it outside the scope of model-risk guidance.
The RBI looked at AI models and wrote them inside.
Same technology. Same quarter. One says not yet. The other says here is how we would like you to think about it.
(A note, since the coverage of the RBI draft has been loud: it is softer than the headlines suggest. It is a consultation draft written in “should,” not “shall” — Guidance, not Directions. It doesn’t mandate kill switches; “override, suspension, or deactivation mechanisms” appear once, as one example in a non-exhaustive list. And it doesn’t ban opaque models — it asks institutions to set explainability thresholds, and expressly allows models that can’t be fully explained provided they carry compensating controls. Worth reading rather than reading about.)
I don’t know which posture is right, and that isn’t false modesty. There’s a real case for each. The Fed’s candour about the limits of a fixed framework is genuine, and Stanley’s point stands — pre-deployment validation of a system that keeps evolving may not be a coherent thing to ask for. But an unspecified standard is one each institution invents privately, and I doubt they all invent it well.
Both positions have a failure mode. Neither is stupid. That’s what makes it worth thinking about rather than merely a story about a regulator getting it wrong.
VII. The sharpest clause is the one nobody is talking about
While the press was writing about a kill switch that isn’t there, they walked straight past this. Paragraph 46, on third-party models:
“…independent validation by the RE… notwithstanding any validation, certification, or assurance provided by the third-party provider.“
And paragraph 45: an institution using a third-party model “is accountable for its outcomes.”
You own the model you didn’t build. A vendor’s safety certificate is not, on this reading, a defence. The draft goes further — contracts should give the institution and its supervisor audit rights and enough technical documentation to validate the model themselves; and where a provider won’t disclose enough, the institution should consider limiting the usage.
Anyone about to buy an enterprise LLM should read that twice. It is the most operationally consequential sentence in either document, and it got no headlines at all, because it doesn’t sound like anything.
(One honest caveat: the draft says “third-party.” It does not say “group entity” or “offshore parent.” Whether a model built by a global parent and pushed into an Indian subsidiary counts as third-party is a reasonable reading — but it is my reading, not RBI’s text.)
VIII. A question I can’t answer from where I sit
Which leads somewhere I can’t resolve.
First, a correction to an instinct I nearly followed. I was going to say: a lot of this AI is built in India’s Global Capability Centres, so those teams sit between two regulators. That’s wrong. A GCC is not a regulated entity — it’s a captive services company. Jurisdiction follows where the institution operates, not where its engineers sit. A model built in Hyderabad and deployed to a US business is a US supervisory matter. Full stop.
But two things drag India back in.
Some of these banks hold RBI banking licences, not just back offices. Citibank N.A. and J.P. Morgan Chase Bank N.A. both do. Those entities are regulated entities, and the draft says so in as many words — “Commercial Banks (including Foreign Banks).”
And the third-party clause has long arms. If a model built centrally is used by the Indian regulated entity, the draft would have that entity independently validate it — whatever anyone concluded in New York.
Which turns my question into a better one:
When your engineering centre and your strictest regulator sit in the same country, what do you actually build to?
Maintaining two versions of an agent platform — one instrumented for the Indian entity, one not — is expensive, and nobody wants to own that fork. The path of least resistance is to build once, to the stricter standard.
If that’s what happens, something quietly strange follows: India’s expectations become the working default for a good deal of what gets built in India — including systems whose home regulator declined to set any.
I want to be clear that this is speculation. I don’t know that it works this way. If you’ve had to reason about it from inside a bank platform team or a model-risk function, I’d like to know how it actually plays out.
IX. What I think it means — held loosely
Here’s my read, as hypothesis rather than finding.
Guidance is not a burden. It’s a permission structure.
When SR 11-7 told you what validation looked like, it also, implicitly, told you what sufficient looked like. You could build to a standard. You could walk into a committee and say: this meets the bar. And the committee, which does not want to be the reason the bank fell behind, could say yes.
Take the standard away and you don’t get freedom. You get a room full of people who can’t tell whether they’re allowed to say yes.
The risk officer then does the only rational thing available: invents a bar. And with nothing to calibrate against — too strict costs you a slow roadmap, too loose costs you your career — the invented bar lands higher than the real one would have.
Which takes me back to the 95%, and the 89%, and the 46%.
If that’s right, then filling the void is a product problem before it’s a compliance problem. Because the artefacts that satisfy a committee turn out to be the same artefacts that make the system good: confidence thresholds with a named owner; rollback that’s been fired in anger rather than documented; audit trails capturing input, model version, score and action; regression suites so “did it get worse?” is a query rather than an argument. And the one nearly everybody skips — a designed, first-class path for the system to say I don’t know. A model with no escape hatch will always answer. That’s not a bug in the model. That’s a missing product decision.
Notice that the RBI’s “override, suspension, deactivation” is a rollback control. Its explainability threshold is an audit trail. The two regimes disagree about what to require. They don’t much disagree about what the artefacts are.
The one thing I’ll say with confidence, because I paid for it: I built this control plane once, in a place where a missed inference meant a frontline worker got hurt rather than a bank got fined. There was no framework; we had to invent one. That’s why I notice this problem — and it’s also, I’m aware, exactly why I might be over-fitting to it.
X. Where I could be wrong
The void might not be real in practice. Large banks may already govern generative AI perfectly well through existing operational-risk and third-party frameworks. From outside, I can’t see that. Someone inside can.
I might be over-fitting to my own scar. I built governance for a system where being wrong hurt someone, so of course I see governance as the binding constraint everywhere. Naming the bias doesn’t dissolve it.
Non-binding guidance might matter less than I think. Maybe the real bar is set by internal politics and three people in a room, in which case my whole argument is aimed at the wrong target.
XI. So
Two regulators looked at the same technology in the same quarter. One decided it was too novel to govern well. The other decided it was too consequential not to try.
I don’t think either of them is being foolish. I think they’ve made opposite bets on the same uncertainty, and the institutions in the middle are going to have to build something before we find out who was right.
Which is, in the end, the answer to the question I started with. AI arrives last where being wrong is expensive — and it arrives there when somebody builds the evidence that lets an institution say yes. Not when the model gets better. The model is already good enough. It arrives when the control plane exists: the thresholds, the rollback, the audit trail, the escape hatch, the named human who owns the number.
That’s not a regulatory artefact. It’s a product.
RBI’s consultation closes on 24 July, and it’s open to anyone. I’m going to write a submission — on override mechanisms as a product primitive rather than a compliance checkbox, and on the limits of demanding an explainability the technology may not yet be able to deliver.
If you live inside this, tell me what I’m getting wrong. Not politely. I’d rather be corrected in public than confident in private.
Srinivas Mullapudi has spent 20 years building enterprise products, data platforms, AI pipelines, and product organizations at the intersection of enterprise software and emerging tech. Ground Truth covers what’s really happening in AI and tech, not what the pitch decks say, with occasional detours into career, life and the messier questions that resist easy frameworks. Check out his portfolio.
Follow along: srinimullapudi.com · Substack · LinkedIn
Sources: SR 26-2 · OCC Bulletin 2026-13 · Chris Stanley, American Banker (16 Jun 2026) · RBI Draft Guidance on Regulatory Principles for Model Risk Management (24 Jun 2026) · MIT NANDA, “The GenAI Divide” (2025) · Gartner and S&P Global Market Intelligence, via public reporting. Bank deployment figures are from public reporting, not filings.






