top of page
vaxa_logo_FF0000_transparent.png
LET'S TALK
vaxa_logo_FF0000_transparent.png
LET'S TALK

Technology

Is AI's ROI Debate Asking the Wrong Question?

Both the bulls and the bears in the AI ROI debate are answering the wrong question. The real fault line is organizational discipline, not which technology camp you're in.

By Vaxa Growth Strategy Team

5 min read

Published on February 18, 2025

Sector Snapshot — Market Intelligence

01

95% of enterprise generative AI pilots deliver zero measurable P&L impact (MIT Project NANDA)

02

>80% overall AI project failure rate — roughly double conventional IT projects (RAND)

03

~6% of organizations attribute 5%+ of EBIT to AI, despite 88% using it somewhere (McKinsey)

04

$6M to train a frontier-competitive model (DeepSeek, Jan 2025) vs. $100M+ for incumbents

Market Consensus & Vaxa's Position

Sector Consenus

AI capex is unprecedented and reckless — a bubble by scale alone.

Vaxa's Position

Wrong on the numbers. By GDP share, this is the same order of magnitude as the railroads. The real risk is concentration in five balance sheets, not size.

Sector Consenus

Task-level gains prove AI works; 95% pilot failure proves it doesn't. Pick a side.

Vaxa's Position

Both camps are answering the wrong question. The real fault line is organizational discipline, not model capability.

Sector Consenus

Proprietary data is the moat once models commoditize.

Vaxa's Position

Half-right. Data without governance discipline is just an asset sitting next to the same failure rate everyone else has.

Sector Consenus

AI needs to find its "killer app" before this spending is justified.

Vaxa's Position

Wrong milestone. Railroads' killer app wasn't railroads — it was refrigeration, built by someone not in the railroad business at all.

Sector Consenus

Model loyalty is a wasting asset — the frontier reshuffles too fast to bet on one. And a single model for every task is a Mercedes for a milk run: the market is splitting into tiers, and a consolidator layer that routes across models is the business this is about to create.

Vaxa's Position

Model loyalty is a wasting asset — the frontier reshuffles too fast to bet on one. And a single model for every task is a Mercedes for a milk run: the market is splitting into tiers, and a consolidator layer that routes across models is the business this is about to create.

Get this wrong and you're either the company that pulled back too early, or the one still defending a model-exclusivity bet nobody asked for.

Historical Precedent

A. The capex isn't reckless. It's concentrated — which is the actual risk.

U.S. railroad investment ran roughly 6–10% of GDP at its 19th-century peak — some estimates put the peak mania years at 10–20% — financed mostly through debt and land grants, with about a third of proposed British lines never built at all. AI capital formation today runs an estimated 1.5–5% of GDP depending on what's counted, and AI-related investment reportedly drove around 74–75% of U.S. GDP growth in the first quarter of 2026 on its own. Same order of magnitude as the railroads. Far more concentrated ownership.

006-chart-A-historical-precedent).png
At Stake

The "unprecedented bubble" framing is wrong on the numbers. The real story isn't scale — it's concentration. Railroad capital was spread across thousands of speculators and hundreds of companies. AI capex sits inside five balance sheets. That's a systemic fragility argument, not a bubble-size argument, and it's a sharper, less-made point than the one dominating the conversation right now.

Technology vs. Application

B. Railroads' killer app wasn't railroads. It was refrigeration.

Refrigerated rail cars weren't part of the original railroad business case. Railroads were built to move freight and passengers faster than existing routes allowed. The refrigerated car — and the meatpacking empire built on top of it — arrived as an afterthought, from someone who wasn't in the railroad business at all. That's the sharper version of the electricity-and-refrigeration parallel often used for this argument: the same infrastructure, built for one purpose, enabling a completely unrelated industry to be built on top of it.

006-chart-A-historical-precedent).png
At Stake

Waiting for AI's obvious "killer app" before committing further is the wrong strategy — it's a strategy that would have missed Armour entirely. The businesses positioned to win aren't the ones waiting for someone to hand them the application. They're the ones with the infrastructure and the intent to go build one.

The Productivity Lag

C. Both sides of the ROI debate are answering the wrong question.

MIT's Project NANDA found that 95% of enterprise generative AI pilots deliver zero measurable P&L impact. RAND separately puts overall AI project failure above 80% — roughly double the failure rate of conventional IT projects. McKinsey's own State of AI survey of nearly 2,000 organizations found that while 88% use AI in at least one business function, only about 6% attribute 5% or more of EBIT to it. Yet task-level gains are real: 14–55% depending on the task, across customer service, coding, and consulting work measured independently. The gap isn't AI failing to work — it's gains not surviving the jump from one person's workflow to the whole company's output, which is the same lag electricity went through when factories bolted motors onto old floor plans instead of redesigning around them.

006-chart-A-historical-precedent).png
At Stake

Both the bulls citing GDP contribution and the bears citing pilot failure are answering the wrong question. MIT's own data says the failure is organizational — roughly 80% of the pilot-to-production gap is data engineering, governance, and workflow integration, not model capability. The real fault line isn't AI skeptics versus AI bulls. It's companies with governance discipline versus companies without it, and that line has nothing to do with which model they bought.

Every new fab represents demand for the same handful of critical equipment suppliers.

Cost Collapse & Roll-Up Economics

D. The frontier reshuffles faster than any exclusivity bet can hold.

DeepSeek trained a frontier-competitive model for roughly $6M in January 2025, against incumbents that had spent $100M or more. Eighteen months later, the field is crowded rather than settled: DeepSeek V4, Qwen 3.6, Kimi K3, and GLM-5 trade the benchmark lead depending on the task, most of them open-weight, several priced under $0.30 per million tokens.

The first business-model disruption: the consolidator. When the underlying commodity gets this cheap this fast, building a front-end that routes each task to the best or cheapest available model becomes the rational business — not building or owning a model at all. Online travel followed the same arc: a handful of owners now run most of the booking platforms travelers think of as independent brands, keeping the storefronts distinct while consolidating the supply and the economics behind them. The same shift is available to whoever builds the routing layer for AI: one front end, many models underneath, priced and selected per task rather than committed to a single vendor.

The second business-model disruption: one size doesn't fit all. A frontier reasoning model priced for the hardest problems is the wrong tool for most of what a business actually needs to do — the same way a Mercedes is the wrong tool for a milk run. Both get the job done; only one is worth the premium for that specific trip. Companies defaulting every task to their most expensive model are paying luxury-car costs for commuter-car problems.

The model half-life is compressing. From 2023 into mid-2025, frontier labs shipped on a roughly six-month cadence per lab. That broke in the second half of 2025 and collapsed in early 2026: the fastest-moving labs are now running four-to-six-week release cadences on their flagship lines.

006-chart-A-historical-precedent).png
At Stake

For most business use cases, model selection is becoming as strategically irrelevant as engine displacement is to a daily commute. The DeepSeek-to-Kimi-to-Qwen-to-GLM sequence in eighteen months proves the frontier reshuffles too fast for any single-model bet to function as a moat.

Where Advantage Survives

E. Data is only a moat with governance discipline attached to it.

The standard answer — proprietary data becomes the moat once models commoditize — is repeated across the industry almost uncritically. McKinsey itself illustrates the tension well: the firm is running 25,000 AI agents alongside its 40,000 human consultants, targeting full parity by year-end, while its own published research finds only about 6% of organizations get real EBIT impact from their AI investment.

One correction worth making precisely: "open source" AI models mostly aren't. What DeepSeek, Qwen, and Kimi actually release is the trained weights and a license to run and fine-tune them — not the training code, not the training data. The more accurate industry term is open-weight. What open-weight deployment actually buys a business isn't transparency into the model — it's privacy for the business's own data.

At Stake

Half-right, and worth saying plainly: data without the governance and integration discipline from Section C isn't a moat — it's an asset sitting next to the same 95% failure rate as everyone else's. The actual moat is the operating discipline to turn data into a measured outcome.

006-chart-A-historical-precedent).png

Bringint it together

01

The capex isn't reckless — it's concentrated, which is a different and more specific risk.

02

Waiting for the obvious application is the wrong strategy, historically speaking.

03

The real fault line in the ROI debate is organizational discipline, not which technology camp you're in.

04

Model loyalty is a wasting asset, and one-size-fits-all model usage is a business-model liability.

05

Which means the only durable moat left is operating discipline applied to data — not the data itself, and not the model.

CLOSING QUESTION

The businesses winning this cycle won't be the ones with the best model or the most data. They'll be the ones who built the discipline to turn either into a measured result. Which one is your organization actually building?

References

Pereira, "Railroads and Economic Growth in the Antebellum United States" (William & Mary working paper)
Magoon, "The Largest Investment Booms in History"
36kr, "Today's AI Infrastructure Building Frenzy: Parallels to the Railway Building Frenzy 150 Years Ago"
BetaFinch, "AI Capex as a Percentage of US GDP in 2026" and "US GDP Q1 2026 Without AI" (2026)
MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025"
RAND Corporation, AI project failure research (2026)
McKinsey Global State of AI Survey (nearly 2,000 organizations, 105 countries, 2025)
Original DeepSeek cost reporting (January 2025)
Remote OpenClaw and Layer3Labs, Chinese AI model comparisons (2026)
McKinsey CEO Bob Sternfels, CES 2026 remarks, reported via The Next Web (May 2026)
Frontier Model Release Velocity Index, Q2 2026 report (digitalapplied.com)
Job Security Meter, "Frontier AI Model Releases 2026: Timeline" (2026)
Open Source Initiative, "Open Weights: Not Quite What You've Been Told" (2025-2026)
PBS News / The Conversation, "What's the Difference Between Closed, Open-Source and Open-Weight AI?" (2026)

bottom of page