top of page
vaxa_logo_FF0000_transparent.png
LET'S TALK
vaxa_logo_FF0000_transparent.png
LET'S TALK
Is AI's ROI Debate Asking the Wrong Question?

Technology

Is AI's ROI Debate Asking the Wrong Question?

Both the bulls and the bears in the AI ROI debate are answering the wrong question. The real fault line is organizational discipline, not which technology camp you're in.

By Vaxa Growth Strategy Team

5 min read

Published on February 18, 2025

HIGHLIGHTS

95% of enterprise generative AI pilots deliver zero measurable P&L impact (MIT Project NANDA)

>80% overall AI project failure rate — roughly double conventional IT projects (RAND)

~6% of organizations attribute 5%+ of EBIT to AI, despite 88% using it somewhere (McKinsey)

$6M to train a frontier-competitive model (DeepSeek, Jan 2025) vs. $100M+ for incumbents

Market Consensus & Vaxa's Position

01. Historical PrecedentSector Consenus

AI capex is unprecedented and reckless — a bubble by scale alone.

Vaxa's Position

Wrong on the numbers. By GDP share, this is the same order of magnitude as the railroads. The real risk is concentration in five balance sheets, not size.

03. The Productivity LagSector Consenus

Task-level gains prove AI works; 95% pilot failure proves it doesn't. Pick a side.

Vaxa's Position

Both camps are answering the wrong question. The real fault line is organizational discipline, not model capability.

05. Where Advantage SurvivesSector Consenus

Proprietary data is the moat once models commoditize.

Vaxa's Position

Half-right. Data without governance discipline is just an asset sitting next to the same failure rate everyone else has.

02. Technology vs. ApplicationSector Consenus

AI capex is unprecedented and reckless — a bubble by scale alone.

Vaxa's Position

Wrong on the numbers. By GDP share, this is the same order of magnitude as the railroads. The real risk is concentration in five balance sheets, not size.

04. Cost Collapse & Roll-Up EconomicsSector Consenus

Being on the frontier model is a competitive moat. One strong model can serve every use case.

Vaxa's Position

Model loyalty is a wasting asset. A single model for every task is a Mercedes for a milk run — the market is splitting into tiers.

At Stake

Get this wrong and you're either the company that pulled back too early, or the one still defending a model-exclusivity bet nobody asked for.

Historical Precedent

A. Is this capex normal, historically?

U.S. railroad investment ran roughly 6–10% of GDP at its 19th-century peak — some estimates put the peak mania years at 10–20% — financed mostly through debt and land grants, with about a third of proposed British lines never built at all. AI capital formation today runs an estimated 1.5–5% of GDP depending on what's counted, and AI-related investment reportedly drove around 74–75% of U.S. GDP growth in the first quarter of 2026 on its own. Same order of magnitude as the railroads. Far more concentrated ownership.

Bar chart comparing railroad capex at its 19th-century peak (~8% of GDP) to AI capex today (~3.25% of GDP)
Vaxa's Position

The "unprecedented bubble" framing is wrong on the numbers. The real story isn't scale — it's concentration. Railroad capital was spread across thousands of speculators and hundreds of companies. AI capex sits inside five balance sheets. That's a systemic fragility argument, not a bubble-size argument, and it's a sharper, less-made point than the one dominating the conversation right now.

Technology vs. Application

B. The refrigeration test

Refrigerated rail cars weren't part of the original railroad business case. Railroads were built to move freight and passengers faster than existing routes allowed. The refrigerated car — and the meatpacking empire built on top of it — arrived as an afterthought, from someone who wasn't in the railroad business at all. That's the sharper version of the electricity-and-refrigeration parallel often used for this argument: the same infrastructure, built for one purpose, enabling a completely unrelated industry to be built on top of it.

insight sec listing hero light 2.png
At Stake

"No one in the medical device industry could solve this. The answer was never going to come from inside medicine. It came from a room with toy designers, psychologists, and technologists — asking a question the industry had never thought to ask."

The Productivity Lag

C. Why the money moves faster than the returns

MIT's Project NANDA found that 95% of enterprise generative AI pilots deliver zero measurable P&L impact. RAND separately puts overall AI project failure above 80% — roughly double the failure rate of conventional IT projects. McKinsey's own State of AI survey of nearly 2,000 organizations found that while 88% use AI in at least one business function, only about 6% attribute 5% or more of EBIT to it. Yet task-level gains are real: 14–55% depending on the task, across customer service, coding, and consulting work measured independently. The gap isn't AI failing to work — it's gains not surviving the jump from one person's workflow to the whole company's output, which is the same lag electricity went through when factories bolted motors onto old floor plans instead of redesigning around them.

006-chart-C-productivity-lag.png
VAXA'S POSITION

Both the bulls citing GDP contribution and the bears citing pilot failure are answering the wrong question. MIT's own data says the failure is organizational — roughly 80% of the pilot-to-production gap is data engineering, governance, and workflow integration, not model capability. The real fault line isn't AI skeptics versus AI bulls. It's companies with governance discipline versus companies without it, and that line has nothing to do with which model they bought.

At Stake

The gap isn't AI failing to work — it's gains not surviving the jump from one person's workflow to the whole company's output.

Cost Collapse & Roll-Up Economics

D. The frontier reshuffles faster than any exclusivity bet can hold

DeepSeek trained a frontier-competitive model for roughly $6M in January 2025, against incumbents that had spent $100M or more. Eighteen months later, the field is crowded rather than settled: DeepSeek V4, Qwen 3.6, Kimi K3, and GLM-5 trade the benchmark lead depending on the task, most of them open-weight, several priced under $0.30 per million tokens.

The first business-model disruption: the consolidator. When the underlying commodity gets this cheap this fast, building a front-end that routes each task to the best or cheapest available model becomes the rational business — not building or owning a model at all. Online travel followed the same arc: a handful of owners now run most of the booking platforms travelers think of as independent brands, keeping the storefronts distinct while consolidating the supply and the economics behind them. The same shift is available to whoever builds the routing layer for AI: one front end, many models underneath, priced and selected per task rather than committed to a single vendor.

The second business-model disruption: one size doesn't fit all. A frontier reasoning model priced for the hardest problems is the wrong tool for most of what a business actually needs to do — the same way a Mercedes is the wrong tool for a milk run. Both get the job done; only one is worth the premium for that specific trip. Companies defaulting every task to their most expensive model are paying luxury-car costs for commuter-car problems. The businesses that segment their model usage by task difficulty — cheap, fast models for routine work, frontier models reserved for what actually requires them — will run at a structurally lower cost basis than competitors who don't. Companies that don't build this segmentation into their strategy default, by omission, into being a single-tier, premium-only shop in a market that's rapidly splitting into tiers.

The model half-life is compressing. From 2023 into mid-2025, frontier labs shipped on a roughly six-month cadence per lab. That broke in the second half of 2025 and collapsed in early 2026: the fastest-moving labs are now running four-to-six-week release cadences on their flagship lines, and the closed labs compressed hard in response — three of the largest collectively shipped seven frontier models in a 78-day window between February and April 2026. Call it the AI-era version of Moore's Law, except the thing halving isn't cost-per-transistor on a fixed clock — it's the useful life of a model-selection decision. A model chosen as "the best available" in January can be three generations behind by year-end.

Bar chart comparing DeepSeek's roughly $6 million training cost against incumbent frontier models at $100 million or more
At Stake

For most business use cases, model selection is becoming as strategically irrelevant as engine displacement is to a daily commute. The DeepSeek-to-Kimi-to-Qwen-to-GLM sequence in eighteen months proves the frontier reshuffles too fast for any single-model bet to function as a moat. Companies that have built their AI strategy around exclusivity with one model are optimizing for a variable that's about to stop mattering — and companies that haven't built tiering into their cost structure at all are quietly choosing to compete only in the Mercedes segment, whether they meant to or not.

Where Advantage Survives

D. The frontier reshuffles faster than any exclusivity bet can hold

The standard answer — proprietary data becomes the moat once models commoditize — is repeated across the industry almost uncritically. McKinsey itself illustrates the tension well: the firm is running 25,000 AI agents alongside its 40,000 human consultants, targeting full parity by year-end, while its own published research finds only about 6% of organizations get real EBIT impact from their AI investment. Data-rich, well-resourced organizations are not automatically exempt from the productivity lag described in Section C.

One correction worth making precisely: "open source" AI models mostly aren't. Traditional open source means the full recipe is public — the code, the training data, the ability to inspect and rebuild the system from scratch. What DeepSeek, Qwen, and Kimi actually release is the trained weights and a license to run and fine-tune them — not the training code, not the training data, no real ability to audit how the model learned what it learned. The more accurate industry term, increasingly preferred over "open source" for exactly this reason, is open-weight. What open-weight deployment actually buys a business isn't transparency into the model — it's privacy for the business's own data. Self-hosting an open-weight model means prompts, fine-tuning data, and inference traffic never leave infrastructure the company controls, because nothing is being sent to a third party's API. That is the real enterprise case for open-weight adoption, and it's a data-governance argument, not a code-transparency one.

Bar chart comparing DeepSeek's roughly $6 million training cost against incumbent frontier models at $100 million or more
At Stake

Half-right, and worth saying plainly: data without the governance and integration discipline from Section C isn't a moat — it's an asset sitting next to the same 95% failure rate as everyone else's. The actual moat is the operating discipline to turn data into a measured outcome, which is a much smaller, much less comfortable group of companies than "everyone with good data."

Bringing It Together

01

The capex isn't reckless — it's concentrated, a different and more specific risk.

02

Waiting for the obvious application is the wrong strategy, historically speaking.

03

The real fault line is organizational discipline, not which technology camp you're in.

04

Model loyalty is a wasting asset — whoever builds the routing layer across tiers captures the consolidator's share.

05

The only durable moat left is operating discipline applied to data — not the data itself, and not the model.

closing question

The businesses winning this cycle won't be the ones with the best model or the most data. They'll be the ones who built the discipline to turn either into a measured result. Which one is your organization actually building?

Vaxa helps leadership teams cut through the AI ROI debate with a clear-eyed read on where the real exposure and the real advantage sit.

References

Pereira, "Railroads and Economic Growth in the Antebellum United States" (William & Mary)

bottom of page