Every firm runs two model factories. Only one of them is on the roadmap.
A few weeks ago I saw the Bardess Data Science Maturity Curve for the first time. It’s a great piece of work — five stages, Learning through Cultural, plotted along an S-curve, with the payoff of each stage annotated above the line and the friction below it. If you’ve spent any time watching an analytics function grow up, you’ll recognize your own organization somewhere on it. Special thanks to Michel Nahon for putting it in front of me, and the subsequent great conversations that led to this post.
I read the piece twice. The first time reminded me of the department within my former employer: 600+ of us, all focused on different parts of that curve. But the second time I wasn’t thinking about data science at all.
I was thinking about the workbooks. The ones that price the deals, size the capital, set the reserves, and produce the number that goes in front of the board. Because the same curve — stage for stage, plateau for plateau, fork for fork — describes how a financial institution matures at financial modeling. And most institutions are years further along one curve than the other without ever having noticed, because only one of the two curves was ever drawn.
Financial Model Maturity measures how far an organization has moved its spreadsheet estate from a collection of private files toward governed, versioned, callable business logic. It runs across five stages — Ad Hoc, Aware, Managed, Operational, Compounding — and it applies to the workbooks that price the deals, size the capital and set the reserves.
It is not the same thing as end-user computing risk management. EUC frameworks exist to contain spreadsheets. This one measures whether they are getting better.
Before going further I should deal with the obvious objection: the world does not need another maturity model. One academic survey catalogs eleven for analytics alone — TDWI, Gartner, DELTA Plus, SAS and others — and nearly all of them descend from the Capability Maturity Model that Carnegie Mellon’s Software Engineering Institute published in 1991. My stage names borrow from it too. Ad Hoc and Managed are CMMI vocabulary, and I’d rather acknowledge the lineage than pretend to have invented a shape that has been in service for thirty-five years.
Two things justify drawing another line anyway.
The first is corroboration. TDWI’s model describes a chasm between early adoption and corporate adoption — obstacles that reliably stall organizations partway up. They arrived at it independently, studying a different discipline, and it sits at the same place on the curve as the Bardess plateau and the one I’m about to describe. Three frameworks finding the same stall in the same position is evidence that the stall is real rather than rhetorical.
The second is the gap. Read down the list and notice what none of them measures.
| Framework | Origin | What it actually measures |
|---|---|---|
| CMMI | Carnegie Mellon SEI, 1991 | Process capability. The ancestor of essentially every five-stage model since, including this one. |
| TDWI Analytics Maturity | TDWI | Analytics adoption — and uniquely, a named chasm between early and corporate adoption. |
| Gartner Data & Analytics | Gartner | Basic → Opportunistic → Systematic → Differentiating → Transformational. |
| DELTA Plus | International Institute for Analytics | Data, Enterprise, Leadership, Targets, Analysts — organizational readiness to compete on analytics. |
| SAS Analytics Maturity | SAS | Unaware → Aware → Astute → Empowered → Explorative. |
| Bardess Data Science Maturity | Bardess | Data science capability, Learning through Cultural. The curve this piece extends. |
| EUC risk frameworks | Regulators and EUC tooling vendors | Spreadsheet risk to be contained — inventoried, locked down, suppressed. Not a maturity curve; it has no stage 5. |
| Financial Model Maturity | This piece | The spreadsheet estate as a capability that can get better, rather than a liability to be minimized. |
Every one of these treats analytics, data or process as a capability that matures. The spreadsheet estate appears in none of them — and where it does appear, in the EUC risk literature and the tooling built around it, it appears exclusively as something to be contained: inventoried, locked down, suppressed, ideally eliminated.
Containment is not a maturity curve. It has no stage 5. It is a strategy for making a liability smaller — a fine thing to do, and a completely different thing from making it an asset.
Nobody has drawn a curve on which the spreadsheet estate gets better. That is the gap this is trying to fill.
The Bardess curve predates the generative AI era, and it would be easy to treat that as a reason to discount it. I’d argue the opposite. The last two years ran an experiment on their framework, and the results are more interesting than the framework alone.
Bardess names the stage-1 friction precisely: talented data scientists are hard to find and expensive. That constraint has softened considerably — a language model will write you a modeling pipeline. The lower stages of the data science curve have partly collapsed since the diagram was drawn.
The spreadsheet curve’s lower stages have not moved at all, because they were never gated on talent. The experiment ran, and it pushed the two curves further apart rather than bringing them together.
This is the part I’d want a stage-3 CFO to sit with. GenAI dramatically increased the rate at which an organization can produce models and analyses. It changed nothing whatsoever about the path from a finished model to a governed production system. So stage-3 organizations now generate more things that cannot ship. Pilot counts went up; the proportion reaching production went down.
You will have seen the widely circulated MIT figure — 95% of surveyed GenAI projects showing no measurable P&L impact. Treat it carefully: the sample was not random, and “no measurable impact” mostly reflects the absence of a pre-deployment baseline rather than technical failure. But the direction is not seriously disputed by anyone who has watched a pilot portfolio, and the mechanism is exactly the one this curve describes. Production capacity rose. Operationalization capacity did not.
Their terminal state is predictive analytics becoming part of everyday language — humans growing fluent in analytics. The current terminal state is closer to the inverse: machines consuming your logic, and calling your models as tools. That is not a lateral extension of their framework into a new domain. It is an extension forward in time, and it is where this piece ends up.
On 17 April 2026 the Federal Reserve issued SR 26-2, superseding SR 11-7 — the guidance every model risk function in the country has been organized around since 2011 — along with SR 21-8. The new letter is principles-based rather than prescriptive, replaces annual review cycles with a risk-based cadence, tightens the definition of a model, and narrows applicability to institutions above $30 billion in assets.
Practitioner readings also report that generative and agentic AI are explicitly carved out. If that reading holds, it is the most consequential sentence in the document for anyone thinking about AI inside a regulated institution: the supervisory framework has declined to treat generative systems as models.
Which means the auditable layer of a bank cannot be the probabilistic one. It has to be the deterministic one — the governed, versioned, inspectable arithmetic.
Hold that thought. It comes back at the end.
The interesting geometry in the Bardess diagram isn’t the rise. Any maturity model rises. It’s the two deformations in the line.
The curve flattens at stage 3. That’s the stage where the team is real, the roles are defined, the use cases are named — and the work is still being done on laptops, integrated to the business ad hoc and fragile, and, in their words, difficult to operationalize.
Then at stage 4 the line forks. A solid branch accelerates upward. A dashed branch peels off and bends down, labeled risk of stagnation and loss of confidence in the data science function.
Those two features carry the whole argument. A discipline can be skilled, staffed, and stalled at the same time. And what decides which branch you take at stage 4 is not the quality of your models.
Run the same five stages against the model portfolio a bank, insurer or asset manager actually bets money on — the spreadsheets — and the shape holds. So does the plateau. So does the fork.
My stage names run parallel to the Bardess ones. The table below sets them side by side — their column, mine, and the constraint both disciplines are actually hitting, which at every stage turns out to be the same constraint.
| Stage | Data Science Maturity | Financial Model Maturity | The shared constraint |
|---|---|---|---|
| STAGE 1Ad Hoc“Learning” | Aware of data science, no in-house expertise. Talent hard to find and expensive; the jargon is a wall. Huge untapped value, and real risk of being left behind. | Spreadsheets everywhere, owned by no one. No inventory — nobody can say how many models the firm runs, or which copy is authoritative. Logic lives in one analyst’s head and their file-naming habits. | The capability is invisible to management. The value is unmeasured — and so is the exposure. |
| STAGE 2Aware“Emerging” | One to three analysts exploring use cases; data science is one responsibility among several. They handle their own IT — inefficient, with security risk. Integration is hard to maintain, impossible to scale. | Someone has named the problem — an audit finding, a bad number, a new EUC policy. Controls are manual: locked tabs, naming conventions, a change log nobody fills in twice. Scenarios are Save As. | Practitioners are doing infrastructure work nobody hired them for, and every control depends on a person remembering. |
| STAGE 3Managed“Functional” | A dedicated team with defined roles and use cases. Still primarily on laptops, doing their own admin. Integration with business workflows is ad hoc and fragile. Difficult to operationalize. | A real function: standards, templates, a tiered inventory, review gates. Documentation exists — as a separate document, already stale. Models still travel as files, so production means a developer rebuilds the logic. Now there are two, and they diverge. | Discipline has outrun infrastructure. Both are one laptop away from production and cannot cross. Both mistake headcount for progress. |
| STAGE 4Operational“Integrated” | Infrastructure finally supports the work; real integration to databases, ERP and BI. Many successful projects. Dashed branch: stagnation, and loss of confidence in the function. | The model becomes a service. Logic addressable and versioned immutably; inputs editable while the formula chain is locked. One model, many interfaces. Scenarios become overlays, not copies. Dashed branch: governance by confiscation. | The tool does not decide the branch. Whether the practitioner keeps authorship does. |
| STAGE 5Compounding“Cultural” | Use cases in every department. Automated ML and citizen tooling. Seamless integration across data science, BI and infrastructure. Results reported at board level. | Every material calculation is a governed, discoverable, callable model. Models compose. New products assemble from existing logic in weeks. Agents invoke the firm’s real arithmetic. Risk becomes a portfolio question. | The capability stops being a project and becomes a supply. The asset compounds, and outlives the people who built it. |
Two rows of that table deserve more than a cell. Stage 3 is where the line flattens and where most serious institutions actually live. Stage 4 is where it forks. And stage 5 needs a section of its own, because the naive reading of it is a disaster.
Stage 3 is where most serious institutions live, and the instinct there is always to add people. More reviewers. More analysts. More controls testing.
It doesn’t move the curve, because capacity was never the constraint.
The constraint is that the logic and its container are the same object. A model that lives in a file can only be governed by governing files — and files are governed by discipline, which is a headcount cost that grows with the portfolio forever.
That’s the same wall the data science team hits at stage 3, expressed in different technology. Their model is fused to a notebook on a laptop; yours is fused to a workbook on a share. Neither can cross into production without being rebuilt by someone else, and the rebuild is where the second, divergent copy of the truth is born.
Bardess is admirably honest about stage 4. Buying the infrastructure is not the same as crossing. Their dashed branch is stagnation and lost confidence in the function.
The financial modeling version of that branch is specific, predictable, and I have watched it happen more than once. It goes like this:
Governance gets achieved by taking the model away from the person who understands the business it describes. Logic changes now require a ticket. The queue is three weeks. The analyst who needs an answer on Thursday exports to Excel on Wednesday. Within two quarters there is a complete stage-1 shadow estate growing underneath a beautifully governed platform — and increasingly, the governed copy is the one nobody uses.
The solid branch requires something narrower than it sounds: the business keeps authorship of the logic, and the platform makes that safe rather than making it somebody else’s job. Lock the formula chain, leave the inputs open, and stop treating the modeler as the risk to be managed.
Here is where I want to be more careful than maturity models usually are, because I described stage 5 as models composing — one model consuming another, and anyone who has lived through a linked-workbook estate should read that sentence with alarm.
We all know what that looks like in Excel. Book A pulls from Book B pulls from Book C. Somebody moves a row. Somebody renames a folder. The “update links?” dialog appears and everyone clicks whichever button makes it go away. Nobody can say what depends on what, and the only way to find out is to open all of it.
So is composition the payoff of stage 5, or the nightmare wearing better clothes?
Both — and the difference isn’t depth. It’s opacity. Every serious codebase on earth is a dense web of imports and nobody loses sleep over it. What makes Excel’s version pathological is four specific properties of how the dependency gets authored.
| Property of the dependency | Linked workbooks (the spider web) | Governed composition (the supply chain) |
|---|---|---|
| What the edge points at | A coordinate: [Book2.xlsx]Sheet1!$D$47. Insert a row and you are silently wrong. | A declared interface. The consumer binds to a named contract, not to a location. |
| Which version it resolves to | Whatever is at that path today. A package manager with no lockfile, re-resolving on every open. | A pinned version. B publishing v9 does not silently change A’s answer; upgrades are explicit and governed. |
| Visibility from the other side | None. Book B has no idea Book A depends on it, so B’s author cannot assess impact before changing it. | Bidirectional. The producer sees every consumer before making a change. |
| Whether the graph is queryable | Emergent. Answering “what breaks if I change this?” means crawling the whole estate. | First-class and queryable — which makes reverse lookup a product recall: this input was wrong, here is everything downstream. |
Fix those four and composition stops being a spider web and becomes a supply chain. Supply chains have known suppliers, version and lot tracking, and — critically — recall capability.
Composition is a stage-5 capability and a stage-3 anti-pattern. If you don’t yet have declared interfaces, pinned versions and a graph you can query in both directions, linking your models together is not progress. It’s manufacturing the nightmare on purpose. Don’t do it yet.
The agentic framing makes all of this urgent rather than academic.
An orchestrator that knows about a firm’s models and can invoke them as deterministic tools is, to my mind, one of the most valuable things in enterprise AI right now. Ask a language model to compute a waterfall and you get a plausible number. Ask a pinned, governed model and you get the number, reproducibly, with a record of which version answered.
But I want to be precise about where the determinism actually lives, because it’s easy to overclaim here.
The models are deterministic. The composition isn’t. If an agent decides at runtime which models to chain and in what order, you haven’t eliminated nondeterminism — you’ve moved it up a layer.
A user’s ugly hardcoded link is at least reproducible. An LLM-selected chain is not, unless you make it so. What makes agent-invoked composition auditable is the same discipline that makes peer composition safe: the resolved graph has to be recorded on every invocation. Model A v7 consumed model B v3, with these inputs, and produced this number, at this time, for this user. Call it a call manifest — the lockfile of the agentic world.
Without it, you have a very fast way to produce numbers you cannot reconstruct, which is the stage-1 failure mode in a much better suit. With it, you can answer the question every regulator is eventually going to ask — which model produced this figure? — exactly rather than probabilistically.
Inherited assumption drift. A model that is perfectly correct in isolation, consumed by six others whose authors never read its assumptions, in a firm where everyone trusts the number precisely because it came from a governed source. Governance creates trust, and trust is exactly what stops people checking. That risk only exists at stage 5, and it deserves naming.
| Dimension | 1 · Ad Hoc | 2 · Aware | 3 · Managed | 4 · Operational | 5 · Compounding |
|---|---|---|---|---|---|
| Ownership | Whoever built it | An informal power modeler | A named function with standards | The business authors; the platform enforces | Firm-wide — logic is a shared asset |
| Where logic lives | A file, on a laptop | Many files, on a share | Controlled files and templates | An addressable, versioned model | A composable portfolio of models |
| Explainability | The author’s memory | A tab nobody updates | A separate document, already stale | The formula explains itself | Explanation is how the model is read |
| Version & audit | Filename suffixes | A change log, sometimes | Review gates and manual sign-off | Immutable history, captured automatically | Continuous, portfolio-wide |
| Scenarios | Save As | Save As, twenty times | A scenario manager few trust | Overlays on a single base model | Simulation across the book |
| Integration | Copy and paste | Export, then import | Fragile macros and refresh chains | APIs, embedded apps, agents | Composed directly into products |
| Infrastructure burden | Unrecognized | Falls on the modeler | Falls on the modeling team | Carried by the platform | Carried by the platform, invisibly |
| Governance posture | None | Policy on paper | Operational — people-powered | Structural — built in, not bolted on | Cultural — assumed |
| Dominant failure mode | A wrong number nobody can trace | An inventory stale the day it ships | A correct model that cannot be shipped | Shadow Excel growing back underneath | Inherited assumption drift |
Maturity is rarely uniform. Most firms sit at different stages on different rows — strong version discipline and no integration path; excellent documentation standards and a documentation burden so heavy the standards are quietly ignored. Read across the row you care about, not down the column you wish you were in.
An analogy that only flatters isn’t worth much. Three real differences.
Stage 1 on the data science curve is gated by scarce, expensive people. The modeling curve isn’t. The talent is already inside your firm, in quantity, sitting in the business units, fluent in the syntax. Nobody needs to be hired or retrained. The bottleneck is purely structural — which means this curve can be climbed considerably faster than the one it resembles.
No firm has legacy machine learning models in every drawer. Every firm already has thousands of spreadsheets, many load-bearing, most undocumented, some running processes whose author left years ago. Financial model maturity is always remediation and migration at once, which makes sequencing — which models move first — the central question rather than an implementation detail.
A firm can sit at stage 4 on data science and stage 2 on financial modeling simultaneously. In financial services that’s close to the norm: production ML pipelines with lineage, feature stores and model risk sign-off — alongside a capital, pricing or reserving model that travels as an email attachment.
That last one is the part I’d sit with. The gap isn’t explained by sophistication, or budget, or talent. It’s explained by attention. The data science curve got named, sponsored and funded. The spreadsheet curve was never drawn, so nobody was ever accountable for a stage on it.
A spreadsheet estate at stage 2 is not static. It grows, it forks, it accumulates dependencies nobody has mapped, and its risk compounds on exactly the same schedule the asset would have.
So the choice at stage 3 isn’t between moving and standing still. It’s between two things compounding, and picking which one.
That’s the whole argument for treating business logic as software capital: governed, versioned, self-documenting, callable, and owned by the people who understand what it means.
Build Assets, Not Liabilities™
Most people who read this place their own firm at stage 3. If that’s you, the useful next question isn’t whether to move. It’s which model moves first. That’s a call I take myself.
Financial Model Maturity measures how far a firm has moved its spreadsheet estate from private files toward governed, versioned, API-callable business logic. It runs across five stages, from Ad Hoc through Compounding, and applies to the workbooks that price deals, size capital and set reserves.
Ad Hoc — spreadsheets owned by no one, no inventory. Aware — the problem is named, controls are manual. Managed — a real function with standards, but models still travel as files. Operational — logic is addressable, versioned and permissioned. Compounding — every material calculation is governed, discoverable and callable.
End-user computing risk management treats spreadsheets as a liability to be contained — inventoried, locked down, ideally eliminated. Financial Model Maturity treats the same estate as a capability that can improve. Containment has no stage 5: it makes a liability smaller rather than turning it into an asset.
Because the logic and its container are the same object. A model that lives in a file can only be governed by governing files, and files are governed by human discipline — a headcount cost that grows with the portfolio forever. Adding reviewers does not move the curve.
SR 26-2, issued 17 April 2026, supersedes SR 11-7 and SR 21-8. It is principles-based, replaces annual review with a risk-based cadence, tightens the definition of a model, and applies to institutions above $30 billion in assets. Practitioner readings also report a carve-out for generative and agentic AI.
Only once the dependency is a declared interface, pinned to a version, visible from both sides, in a graph you can query in both directions. Meet those four conditions and composition is a supply chain. Miss any of them and it is the Excel linked-workbook problem with better branding.