Technology
Is AI's ROI Debate Asking the Wrong Question?
Both the bulls and the bears in the AI ROI debate are answering the wrong question. The real fault line is organizational discipline, not which technology camp you're in.
By Vaxa Growth Strategy Team
5 min read
Published on February 18, 2025
Sector Snapshot — Market Intelligence
01
95% of enterprise generative AI pilots deliver zero measurable P&L impact (MIT Project NANDA)
02
>80% overall AI project failure rate — roughly double conventional IT projects (RAND)
03
~6% of organizations attribute 5%+ of EBIT to AI, despite 88% using it somewhere (McKinsey)
04
$6M to train a frontier-competitive model (DeepSeek, Jan 2025) vs. $100M+ for incumbents
Market Consensus & Vaxa's Position
AI capex is unprecedented and reckless — a bubble by scale alone.
Wrong on the numbers. By GDP share, this is the same order of magnitude as the railroads. The real risk is concentration in five balance sheets, not size.
AI needs to find its "killer app" before this spending is justified.
Wrong milestone. Railroads' killer app wasn't railroads — it was refrigeration, built by someone not in the railroad business at all.
Model loyalty is a wasting asset — the frontier reshuffles too fast to bet on one. And a single model for every task is a Mercedes for a milk run: the market is splitting into tiers, and a consolidator layer that routes across models is the business this is about to create.
Model loyalty is a wasting asset — the frontier reshuffles too fast to bet on one. And a single model for every task is a Mercedes for a milk run: the market is splitting into tiers, and a consolidator layer that routes across models is the business this is about to create.
Get this wrong and you're either the company that pulled back too early, or the one still defending a model-exclusivity bet nobody asked for.
Historical Precedent
A. The capex isn't reckless. It's concentrated — which is the actual risk.
U.S. railroad investment ran roughly 6–10% of GDP at its 19th-century peak — some estimates put the peak mania years at 10–20% — financed mostly through debt and land grants, with about a third of proposed British lines never built at all. AI capital formation today runs an estimated 1.5–5% of GDP depending on what's counted, and AI-related investment reportedly drove around 74–75% of U.S. GDP growth in the first quarter of 2026 on its own. Same order of magnitude as the railroads. Far more concentrated ownership.
.png)
The "unprecedented bubble" framing is wrong on the numbers. The real story isn't scale — it's concentration. Railroad capital was spread across thousands of speculators and hundreds of companies. AI capex sits inside five balance sheets. That's a systemic fragility argument, not a bubble-size argument, and it's a sharper, less-made point than the one dominating the conversation right now.
Technology vs. Application
B. Railroads' killer app wasn't railroads. It was refrigeration.
Refrigerated rail cars weren't part of the original railroad business case. Railroads were built to move freight and passengers faster than existing routes allowed. The refrigerated car — and the meatpacking empire built on top of it — arrived as an afterthought, from someone who wasn't in the railroad business at all. That's the sharper version of the electricity-and-refrigeration parallel often used for this argument: the same infrastructure, built for one purpose, enabling a completely unrelated industry to be built on top of it.

Waiting for AI's obvious "killer app" before committing further is the wrong strategy — it's a strategy that would have missed Armour entirely. The businesses positioned to win aren't the ones waiting for someone to hand them the application. They're the ones with the infrastructure and the intent to go build one.
The Productivity Lag
C. Both sides of the ROI debate are answering the wrong question.
MIT's Project NANDA found that 95% of enterprise generative AI pilots deliver zero measurable P&L impact. RAND separately puts overall AI project failure above 80% — roughly double the failure rate of conventional IT projects. McKinsey's own State of AI survey of nearly 2,000 organizations found that while 88% use AI in at least one business function, only about 6% attribute 5% or more of EBIT to it. Yet task-level gains are real: 14–55% depending on the task, across customer service, coding, and consulting work measured independently. The gap isn't AI failing to work — it's gains not surviving the jump from one person's workflow to the whole company's output, which is the same lag electricity went through when factories bolted motors onto old floor plans instead of redesigning around them.

Both the bulls citing GDP contribution and the bears citing pilot failure are answering the wrong question. MIT's own data says the failure is organizational — roughly 80% of the pilot-to-production gap is data engineering, governance, and workflow integration, not model capability. The real fault line isn't AI skeptics versus AI bulls. It's companies with governance discipline versus companies without it, and that line has nothing to do with which model they bought.
Every new fab represents demand for the same handful of critical equipment suppliers.
Cost Collapse & Roll-Up Economics
D. The frontier reshuffles faster than any exclusivity bet can hold.
DeepSeek trained a frontier-competitive model for roughly $6M in January 2025, against incumbents that had spent $100M or more. Eighteen months later, the field is crowded rather than settled: DeepSeek V4, Qwen 3.6, Kimi K3, and GLM-5 trade the benchmark lead depending on the task, most of them open-weight, several priced under $0.30 per million tokens.
The first business-model disruption: the consolidator. When the underlying commodity gets this cheap this fast, building a front-end that routes each task to the best or cheapest available model becomes the rational business — not building or owning a model at all. Online travel followed the same arc: a handful of owners now run most of the booking platforms travelers think of as independent brands, keeping the storefronts distinct while consolidating the supply and the economics behind them. The same shift is available to whoever builds the routing layer for AI: one front end, many models underneath, priced and selected per task rather than committed to a single vendor.
The second business-model disruption: one size doesn't fit all. A frontier reasoning model priced for the hardest problems is the wrong tool for most of what a business actually needs to do — the same way a Mercedes is the wrong tool for a milk run. Both get the job done; only one is worth the premium for that specific trip. Companies defaulting every task to their most expensive model are paying luxury-car costs for commuter-car problems.
The model half-life is compressing. From 2023 into mid-2025, frontier labs shipped on a roughly six-month cadence per lab. That broke in the second half of 2025 and collapsed in early 2026: the fastest-moving labs are now running four-to-six-week release cadences on their flagship lines.

Where Advantage Survives
E. Data is only a moat with governance discipline attached to it.
The standard answer — proprietary data becomes the moat once models commoditize — is repeated across the industry almost uncritically. McKinsey itself illustrates the tension well: the firm is running 25,000 AI agents alongside its 40,000 human consultants, targeting full parity by year-end, while its own published research finds only about 6% of organizations get real EBIT impact from their AI investment.
One correction worth making precisely: "open source" AI models mostly aren't. What DeepSeek, Qwen, and Kimi actually release is the trained weights and a license to run and fine-tune them — not the training code, not the training data. The more accurate industry term is open-weight. What open-weight deployment actually buys a business isn't transparency into the model — it's privacy for the business's own data.
Half-right, and worth saying plainly: data without the governance and integration discipline from Section C isn't a moat — it's an asset sitting next to the same 95% failure rate as everyone else's. The actual moat is the operating discipline to turn data into a measured outcome.
.png)
Bringint it together
01
The capex isn't reckless — it's concentrated, which is a different and more specific risk.
02
Waiting for the obvious application is the wrong strategy, historically speaking.
03
The real fault line in the ROI debate is organizational discipline, not which technology camp you're in.
04
Model loyalty is a wasting asset, and one-size-fits-all model usage is a business-model liability.
05
