Technology
Is AI's ROI Debate Asking the Wrong Question?
Both the bulls and the bears in the AI ROI debate are answering the wrong question. The real fault line is organizational discipline, not which technology camp you're in.
By Vaxa Growth Strategy Team
5 min read
Published on February 18, 2025
HIGHLIGHTS
95% of enterprise generative AI pilots deliver zero measurable P&L impact (MIT Project NANDA)
>80% overall AI project failure rate — roughly double conventional IT projects (RAND)
~6% of organizations attribute 5%+ of EBIT to AI, despite 88% using it somewhere (McKinsey)
$6M to train a frontier-competitive model (DeepSeek, Jan 2025) vs. $100M+ for incumbents
Market Consensus & Vaxa's Position
AI capex is unprecedented and reckless — a bubble by scale alone.
Wrong on the numbers. By GDP share, this is the same order of magnitude as the railroads. The real risk is concentration in five balance sheets, not size.
Historical Precedent
A. Is this capex normal, historically?
U.S. railroad investment ran roughly 6–10% of GDP at its 19th-century peak — some estimates put the peak mania years at 10–20% — financed mostly through debt and land grants, with about a third of proposed British lines never built at all. AI capital formation today runs an estimated 1.5–5% of GDP depending on what's counted, and AI-related investment reportedly drove around 74–75% of U.S. GDP growth in the first quarter of 2026 on its own. Same order of magnitude as the railroads. Far more concentrated ownership.

The "unprecedented bubble" framing is wrong on the numbers. The real story isn't scale — it's concentration. Railroad capital was spread across thousands of speculators and hundreds of companies. AI capex sits inside five balance sheets. That's a systemic fragility argument, not a bubble-size argument, and it's a sharper, less-made point than the one dominating the conversation right now.
Technology vs. Application
B. The refrigeration test
Refrigerated rail cars weren't part of the original railroad business case. Railroads were built to move freight and passengers faster than existing routes allowed. The refrigerated car — and the meatpacking empire built on top of it — arrived as an afterthought, from someone who wasn't in the railroad business at all. That's the sharper version of the electricity-and-refrigeration parallel often used for this argument: the same infrastructure, built for one purpose, enabling a completely unrelated industry to be built on top of it.

The Productivity Lag
C. Why the money moves faster than the returns
MIT's Project NANDA found that 95% of enterprise generative AI pilots deliver zero measurable P&L impact. RAND separately puts overall AI project failure above 80% — roughly double the failure rate of conventional IT projects. McKinsey's own State of AI survey of nearly 2,000 organizations found that while 88% use AI in at least one business function, only about 6% attribute 5% or more of EBIT to it. Yet task-level gains are real: 14–55% depending on the task, across customer service, coding, and consulting work measured independently. The gap isn't AI failing to work — it's gains not surviving the jump from one person's workflow to the whole company's output, which is the same lag electricity went through when factories bolted motors onto old floor plans instead of redesigning around them.

Both the bulls citing GDP contribution and the bears citing pilot failure are answering the wrong question. MIT's own data says the failure is organizational — roughly 80% of the pilot-to-production gap is data engineering, governance, and workflow integration, not model capability. The real fault line isn't AI skeptics versus AI bulls. It's companies with governance discipline versus companies without it, and that line has nothing to do with which model they bought.
Cost Collapse & Roll-Up Economics
D. The frontier reshuffles faster than any exclusivity bet can hold
DeepSeek trained a frontier-competitive model for roughly $6M in January 2025, against incumbents that had spent $100M or more. Eighteen months later, the field is crowded rather than settled: DeepSeek V4, Qwen 3.6, Kimi K3, and GLM-5 trade the benchmark lead depending on the task, most of them open-weight, several priced under $0.30 per million tokens.
The first business-model disruption: the consolidator. When the underlying commodity gets this cheap this fast, building a front-end that routes each task to the best or cheapest available model becomes the rational business — not building or owning a model at all. Online travel followed the same arc: a handful of owners now run most of the booking platforms travelers think of as independent brands, keeping the storefronts distinct while consolidating the supply and the economics behind them. The same shift is available to whoever builds the routing layer for AI: one front end, many models underneath, priced and selected per task rather than committed to a single vendor.
The second business-model disruption: one size doesn't fit all. A frontier reasoning model priced for the hardest problems is the wrong tool for most of what a business actually needs to do — the same way a Mercedes is the wrong tool for a milk run. Both get the job done; only one is worth the premium for that specific trip. Companies defaulting every task to their most expensive model are paying luxury-car costs for commuter-car problems. The businesses that segment their model usage by task difficulty — cheap, fast models for routine work, frontier models reserved for what actually requires them — will run at a structurally lower cost basis than competitors who don't. Companies that don't build this segmentation into their strategy default, by omission, into being a single-tier, premium-only shop in a market that's rapidly splitting into tiers.
The model half-life is compressing. From 2023 into mid-2025, frontier labs shipped on a roughly six-month cadence per lab. That broke in the second half of 2025 and collapsed in early 2026: the fastest-moving labs are now running four-to-six-week release cadences on their flagship lines, and the closed labs compressed hard in response — three of the largest collectively shipped seven frontier models in a 78-day window between February and April 2026. Call it the AI-era version of Moore's Law, except the thing halving isn't cost-per-transistor on a fixed clock — it's the useful life of a model-selection decision. A model chosen as "the best available" in January can be three generations behind by year-end.

For most business use cases, model selection is becoming as strategically irrelevant as engine displacement is to a daily commute. The DeepSeek-to-Kimi-to-Qwen-to-GLM sequence in eighteen months proves the frontier reshuffles too fast for any single-model bet to function as a moat. Companies that have built their AI strategy around exclusivity with one model are optimizing for a variable that's about to stop mattering — and companies that haven't built tiering into their cost structure at all are quietly choosing to compete only in the Mercedes segment, whether they meant to or not.
Where Advantage Survives
D. The frontier reshuffles faster than any exclusivity bet can hold
The standard answer — proprietary data becomes the moat once models commoditize — is repeated across the industry almost uncritically. McKinsey itself illustrates the tension well: the firm is running 25,000 AI agents alongside its 40,000 human consultants, targeting full parity by year-end, while its own published research finds only about 6% of organizations get real EBIT impact from their AI investment. Data-rich, well-resourced organizations are not automatically exempt from the productivity lag described in Section C.
One correction worth making precisely: "open source" AI models mostly aren't. Traditional open source means the full recipe is public — the code, the training data, the ability to inspect and rebuild the system from scratch. What DeepSeek, Qwen, and Kimi actually release is the trained weights and a license to run and fine-tune them — not the training code, not the training data, no real ability to audit how the model learned what it learned. The more accurate industry term, increasingly preferred over "open source" for exactly this reason, is open-weight. What open-weight deployment actually buys a business isn't transparency into the model — it's privacy for the business's own data. Self-hosting an open-weight model means prompts, fine-tuning data, and inference traffic never leave infrastructure the company controls, because nothing is being sent to a third party's API. That is the real enterprise case for open-weight adoption, and it's a data-governance argument, not a code-transparency one.

Half-right, and worth saying plainly: data without the governance and integration discipline from Section C isn't a moat — it's an asset sitting next to the same 95% failure rate as everyone else's. The actual moat is the operating discipline to turn data into a measured outcome, which is a much smaller, much less comfortable group of companies than "everyone with good data."
Bringing It Together
01
The capex isn't reckless — it's concentrated, a different and more specific risk.
02
Waiting for the obvious application is the wrong strategy, historically speaking.
03
The real fault line is organizational discipline, not which technology camp you're in.
04
Model loyalty is a wasting asset — whoever builds the routing layer across tiers captures the consolidator's share.
05
