instrui · briefingsIndependent analyst briefing

The Money Ledger · Cost · Edition One

The 95% Problem Is a Money Problem

Why enterprise AI pilots die broke — and the financial discipline behind the ones that don’t.

Jason A. Milne  ·  Published Aug 9, 2026  ·  Briefing edition  ·  9 min read

Verified against public sources through Aug 9, 2026Public sources only

Independent work: not affiliated with, endorsed, or reviewed by any organization cited. Not investment advice. Prepared with AI assistance; see Method & Provenance.

Bottom line up front

The famous 95% measures pilots without P&L impact — a financial finding, not a technological one, and 2026 CEO data confirms the direction.

AI spend went from 31% to 98% of FinOps practice in two years, while roughly 73% of agentic programs blew budget: the control system is being retrofitted mid-flight.

The 5% share four habits, all financial: partner over build, target the back office, integrate with learning loops, and measure before deploying.

Tokens behave like COGS, not software licenses — and beneath them sits the physical layer, GPUs, power, and facilities, now the industry’s binding constraint.

The market’s answer, forward deployment, is a multibillion-dollar bet with no controlled outcome evidence yet. Buyers should contract for the graduation metrics themselves.

§ 01 · The number, honestly

What the 95% actually measures

Everyone in enterprise AI can recite the number. Almost nobody can tell you what it measures.

The number is 95%, and it comes from MIT NANDA’s GenAI Divide: State of AI in Business 2025 report: despite an estimated $30–40 billion in enterprise GenAI investment, 95% of corporate pilots showed no measurable P&L impact. The research drew on 52 structured interviews, 153 survey responses from senior leaders, and a systematic review of more than 300 publicly disclosed AI initiatives.

Here is what the number does not say: that the models don’t work, or that AI is hype. The study measured whether pilots produced P&L impact — a narrow, financial question — and its methodology drew fair criticism after the headline went viral. Both things can be true: the stat is imperfect, and the direction is right. PwC’s January 2026 CEO survey found 56% of chief executives saw neither revenue growth nor cost reduction from AI in the prior twelve months. Only about one in eight — PwC’s “vanguard” — achieved both. Across every serious 2026 dataset, substantial ROI at scale remains the exception.

So the interesting question was never “does AI work?” It’s why the money keeps disappearing between the pilot and the P&L.

§ 02 · The misdiagnosis

It was never the models

MIT’s own authors are explicit about this. The failure driver isn’t model quality or regulation — it’s what they call a learning gap: tools that can’t retain feedback or adapt to context, inside organizations that never wired them into real workflows.

The strangest evidence sits in the same report. While official pilots stalled, roughly 90% of surveyed workers were using personal AI tools daily — against only about 40% of firms holding official subscriptions. A shadow AI economy was producing value at the individual level while the sanctioned programs produced slideware. Adoption was never the constraint. The economics of transformation were.

I spent the early part of my career inside Intel’s D1X fab, labeling nanoscale defect imagery that fed yield models across more than forty metrology tools. Here’s the thing about yield: when it slipped, it was never because the physics stopped working. It was because measurement discipline had — a drifting tool, an unattributed variable, a feedback loop nobody closed. Enterprise AI in 2026 is failing the same way. The capability is real. The control system around it is missing.

§ 03 · Follow the money

The budget carnage, itemized

If you want to see the control system being retrofitted in real time, look at the FinOps community. The FinOps Foundation’s State of FinOps 2026 reports that 98% of practitioners now manage AI spend — up from 63% in 2025 and just 31% in 2024. No cost domain has ever been absorbed into the discipline that fast, and AI cost management now tops the list of skills teams say they need.

They’re absorbing it because the budgets are on fire. One 2026 review of 127 enterprise agentic AI implementations found 73% ran over budget — some by more than 2.4x, burning millions in costs nobody forecast.

Cost ledger · Why AI budgets break
DriverMechanism
Token pricingCost varies request by request with prompt length, output length, model tier, and caching behavior; practitioner reports out of FinOps X describe monthly AI spend swinging as much as 300% on usage patterns alone.
Agentic multipliersAn agent that plans, retries, and calls tools can consume an order of magnitude more tokens than the single API call someone budgeted.
Shared GPU estatesClusters running mixed training and inference have no clean cost boundaries, so attribution fails before optimization can start.
The invisible majorityThe model bill — the line everyone stares at — is rarely the biggest cost in the program; data readiness, integration, and adoption eat the rest, mostly unmetered.

Judgments are this briefing’s; sources in appendix.

And the stack runs deeper than tokens. Under every inference call sits a physical layer — GPUs, cooling, switchgear, megawatts — that most AI budget conversations treat as someone else’s line item. I’ve spent a good part of my career designing and delivering the electrical and mechanical systems buildings actually run on, and the AI buildout has turned that unglamorous layer into the industry’s binding constraint. A cost discipline that stops at the API bill isn’t a cost discipline; it’s a partial view of one.

The 95% didn’t fail because AI was too expensive. They failed because nobody could say what anything cost, per unit of work, until the money was gone.

§ 04 · The survivors

What the 5% actually do

Strip the case studies down and the survivors share four habits, all boring, all financial:

They buy or partner instead of building. MIT found externally sourced tools succeeded roughly twice as often as internal builds — not because vendors are smarter, but because partnerships arrive with deployment discipline attached.

They aim at the back office. Enterprise AI budgets flow overwhelmingly to sales and marketing; the measurable ROI concentrates in operations and finance, where processes have countable baselines.

They integrate into workflows with memory and learning loops. Generic chat tools win adoption on trivial tasks and stall the moment work demands context. The systems that scale are embedded, remember, and improve.

They design measurement in before deployment. The successful programs baseline the process first — labor hours at loaded cost, cycle time, error and rework rates — so “did it work” is an arithmetic question, not a vibes question. And they fund the human side: implementation guides consistently put change management at 15–20% of project budget, and it is reliably the first line item zeroed out.

None of this is machine learning. All of it is management accounting.

§ 05 · The atomic unit

Tokens are the new unit economics

Which brings us to the atomic unit. The correct question about an AI system is not “what does it cost per month?” It’s “what does it cost per unit of work — per resolved ticket, per processed document, per completed task — and is that number going down while quality holds?”

I feel this at small scale in my own lab. Every evaluation suite I run against an LLM pipeline has a token bill attached; every architecture choice — model tier, prompt design, caching, retrieval depth — moves the cost per task. That is COGS behavior, not software-license behavior, and it demands COGS instruments: telemetry joining tokens and GPU time to features and teams, showback, then chargeback, then budget gates tied to value.

The institutions are catching up fast. The FinOps Foundation shipped its FinOps for AI certification — the first formal credential for AI spend management — with the exam going live in March 2026. The Linux Foundation stood up a dedicated Tokenomics Foundation. And this September, the first Tokenomicon conference runs alongside FinOps X in Amsterdam, built on a framing I’d paraphrase this way: FinOps asks whether the AI bill matches the plan; tokenomics asks whether the plan can get cheaper without losing quality. Most organizations have staffed one of those questions and assumed the other would answer itself.

§ 06 · The market’s answer

Send in the engineers

The industry has noticed where the bodies are buried, and its response has been to throw humans at the gap. The forward-deployed engineer — the Palantir-pioneered role that embeds builders inside customer environments — has become the hottest hire in AI. Indeed data shows FDE postings up 729% year over year. Executive search firm Christian & Timbers estimates only about 2,000 engineers in the entire U.S. combine the applied-AI skill and sector fluency to reliably deliver enterprise AI ROI — against demand it projects will surge twenty-fold, with large consultancies planning to grow FDE teams by 10x.

The capital is following. AWS launched a Forward Deployed Engineering organization backed by $1 billion, embedding pods of engineers with customers on ~45-day engagements. OpenAI stood up a dedicated Deployment company with over $4 billion behind it, acquiring the consultancy Tomoro and its 150 forward-deployed engineers outright. Databricks formalized an FDE organization of its own, citing work with more than 1,900 customers over the preceding twelve months.

There’s a live argument about whether this is genius or regression. Andreessen Horowitz frames it as trading margin for moat. Skeptics see a services trap — every customer a fresh consulting project. The sharpest critique asks whether forward deployment builds a more capable customer or a more dependent one.

§ 07 · The evidence gap

The part nobody says out loud

There is, as of today, no controlled evidence that FDE-led deployments beat the 95% base rate.

The evidence is market behavior and vendor self-reporting — billions in bets, exploding job postings, launch-blog case studies — and none of it is controlled outcome research. The bet is plausible; MIT’s data shows partnered deployments succeed more often, and FDEs are partnership in its most concentrated form. But plausible is not proven. What would settle it is unglamorous: pilot-to-production graduation rates, time-to-P&L, and cost-per-outcome for FDE-led deployments against matched programs without them. Until someone publishes that, forward deployment is a rational hedge, not a cure — and buyers should contract for the graduation metrics themselves.

§ 08 · The playbook

Crawl, walk, run

If you own an AI program’s economics, the sequence is not mysterious. It’s process control.

Crawl: inventory every AI vendor and model in use, sanctioned or not. Turn on provider cost APIs. Tag every dollar to a team and a use case.

Walk: run pilots with real token telemetry from day one. Compute cost per unit of work. Publish showback so owners see their own burn.

Run: chargeback. Budget gates tied to value metrics, not calendar months. Evaluation and cost regression in the deployment pipeline — a model change that doubles cost per task should fail the build the same way an accuracy drop does. And fund adoption like the 15–20% line item it is.

Measure, attribute, correct. Any fab operator would recognize it.

§ 09 · So what

The divide is financial

The GenAI Divide was never really about model capability. The 5% treat tokens like COGS and deployments like capital projects — baselined, instrumented, governed. The 95% treat them like software licenses: paid for, rolled out, and assumed.

The gap between those two postures is now a discipline, with a certification, a foundation, and — as of this September — its own conference. This piece is the money ledger.

Jason Milne is a licensed Professional Engineer and data scientist working the full AI cost stack — from the physical infrastructure layer to the digital application layer. He has delivered technical programs and solutions across semiconductor manufacturing, healthcare facilities, and MEP engineering, and builds production LLM systems — multi-agent orchestration, RAG, and evaluation frameworks.

Appendix · Principal sources

  1. MIT NANDA, The GenAI Divide: State of AI in Business 2025 (and Fortune’s interview with lead author Aditya Challapally)
  2. PwC CEO Survey, January 2026
  3. FinOps Foundation, State of FinOps 2026; FinOps for AI certification
  4. 2026 review of 127 enterprise agentic implementations, as reported in FinOps trade coverage
  5. FinOps X 2026 practitioner takeaways on AI cost volatility
  6. TechCrunch (July 2026) on the Christian & Timbers FDE study; Indeed job-postings data on FDE growth
  7. AWS Partner Network announcement of Forward Deployed Engineering
  8. Databricks, “Forward Deployed Engineering: Delivering Business Outcomes with AI,” Jun 2026
  9. TechTarget on OpenAI’s Deployment company and the Tomoro acquisition
  10. Andreessen Horowitz, “Services-Led Growth”; Legion Intelligence, “The Forward Deployed Engineering Model Is Backward”
  11. Linux Foundation / Tokenomics Foundation and Tokenomicon Amsterdam announcements

Method & provenance

First edition, August 2026, of The Money Ledger — a sibling briefing to AI Fault Lines (instrui.ai/posts), applying the same discipline to AI deployment economics. Every statistic is attributed to a named source in-text; source claims are paraphrased throughout; figures were gathered and re-verified against the public record through August 9, 2026. Disclosure: this edition was prepared with the assistance of Claude (an Anthropic model); Anthropic is among the model vendors whose pricing and products are discussed, and readers should weigh that accordingly. Not affiliated with any organization cited. Not investment advice.