← Back to Overview
PUBLICATION TIMESTAMP
--

Enterprise AI Has a Cost Crisis, Not a Model Crisis

Enterprise AI Has a Cost Crisis, Not a Model Crisis

The bills landed before the pilots did. In 2024, enterprises spent an estimated $30 billion to $40 billion on generative AI experiments, according to MIT NANDA research, and 95% of those pilots produced no measurable profit-and-loss impact. Dollar-weighted, that was a rough quarter for the future of enterprise software. Yet the spending curve never bent. Gartner forecasts worldwide AI spending will hit roughly $2.59 trillion in 2026, a 47% year-over-year jump. Enterprise generative-AI spend alone has gone from an estimated $11.5 billion in 2024 to $37 billion in 2025. IDC projects enterprise AI investment rising from about $307 billion in 2025 to $632 billion by 2028. If the top-line numbers look like a boom, the internal numbers look like a different beast: budgets that do not hold, finance teams that do not fully understand the bills, and a widening gap between how fast AI is deployed and how fast its value shows up. | Metric | 2024 | 2025 | 2026 forecast | 2028 forecast | |---|---|---|---|---| | Global AI spending | — | ~$1.5T | $2.52–2.59T | — | | Enterprise generative-AI spend | $11.5B | $37B | — | — | | Enterprise AI investment | — | ~$307B | — | ~$632B | | Average large-company AI budget | $1.2M | — | $7M | — | Source data as reported by Gartner, IDC, and Menlo Ventures. The most worrying number is not the top line. It is the overrun rate. SpendHound’s 2026 AI Spend Report found that 46% of enterprises blew past their AI budgets in 2025. Traditional software overruns ran at 37%. As the report puts it: “The gap is modest but meaningful: AI spend is harder to predict than any other budget line.” Around 81% of finance leaders expect AI budgets to keep increasing in 2026, but 57% are not confident they are paying a fair price. More importantly, 22% say no single person owns the AI budget. That is not a cost-management problem; that is a cost-accounting vacuum. SpendHound also found a striking correlation that doesn’t fit neatly into the usual model-vs-model debate: 62% of Claude users said they had over-budget AI spend, compared with 20% of non-Claude users. Among single-tool users, the gap grew to 73% for Claude-only shops versus 18% for everyone else. This is not an indictment of a particular model. It is a proxy for how enterprises adopt AI. Teams that use deeper, more autonomous tooling are getting more power and more volatility. The same adoption curve that produces better code also produces harder-to-predict bills.

If the spend side is overheating, the return side is flatlining. Domino Data Lab’s fifth annual enterprise AI survey, covering 639 senior AI leaders, found that the share of companies whose ROI fails to outpace investment has held at 57% since 2025. That is a two-year plateau, not a temporary dip. Even when deployment metrics improve — 93% of organizations reported better production capability in 2026, up from 88% in 2025 — the financial outcome does not move. Domino COO Thomas Robinson put the problem in operational terms: “Getting a model into production used to be the milestone that mattered. Our research shows that’s not enough anymore. The real milestone is the moment a business user can act on what the model found, and for too many enterprises, that moment still isn’t happening at the pace or scale of business.” The data backs the quote. 34% of organizations report a mix of AI access methods that varies by business unit, and 40% still rely on at least one mediated access method — a scheduled report from the data science team, or a request to an analyst who runs the analysis and returns the results later. The model may be in production; the business user is still waiting for an email attachment. McKinsey’s well-known “80/80” problem describes the same phenomenon from a different angle: 80% of companies report using the latest generation of AI, and 80% report no meaningful top-line or bottom-line gains. Horizontal assistants that save individual employees a few minutes here and there do not automatically become enterprise value. The saved time is fragmented, unmeasured, and rarely reinvested into a process someone actually owns. Not every survey is doom. Google’s 2025 study of 2,508 large-company executives found that 74% reported positive ROI in at least one generative AI use case, usually within a year. But the same study found that only a small group of “leaders” — roughly 16% of respondents — were capturing outsized revenue gains. BCG’s global survey of 1,250 companies is blunter: only 5% of companies have achieved AI value at scale, while 60% have invested without meaningful return. The gap between “some pilots work” and “the company’s economics change” is the real enterprise AI story.

Gartner’s 2028 forecast was never only about the future

Gartner’s “10 Best Practices for Optimizing Generative and Agentic AI Costs” report contains a prediction that is increasingly being read as a description of the present: “Through 2028, at least 50% of GenAI projects will overrun their budgeted costs due to poor architectural choices and lack of operational know-how.” Gartner has also flagged the structural reason: “Creating a production-ready GenAI system can be orders of magnitude more expensive than running a pilot.” Pilots are forgiving. Production systems have latency, security, data, and concurrency constraints that multiply token consumption and infrastructure spend. Inference, not training, is the hidden cost driver. Training is a large, visible upfront expense. Inference is recurring and sneaky — it happens on every user prompt, every automated workflow, every retry. Gartner expects inference to be at least 70% of a model’s lifetime costs. For enterprises that planned around a training budget and then discovered the inference meter, the 70% number feels less like a forecast and more like a confession. The same Gartner analysis turns AI coding into an accounting problem. By 2028, AI coding costs will overtake the average developer’s salary, driven by rising LLM token consumption and the shift to consumption-based licensing. The billing model alone makes the cost unpredictable. If a developer’s tool license can multiply as they use more tokens, then a company’s software engineering budget is no longer tied to headcount; it is tied to every autocomplete and every agent execution. Gartner notes that lack of transparency from vendors increases the risk of overruns as organizations scale AI-assisted development. Agentic AI takes this to another level. An agent doesn’t just generate a response; it calls models repeatedly, retrieves files, invokes tools, and executes multi-step workflows. Traditional chatbots have a fixed cost per query. Agents have a cost per journey. Gartner expects more than 40% of agentic AI projects to be canceled by the end of 2027 because of escalating costs, unclear business value, and inadequate risk controls. That is not a skeptical aside; it is the largest analyst firm on the planet saying that most agentic initiatives will die in spreadsheets before they survive in production.

Why enterprises can’t get a grip on the cost

Suplari’s finance-leader research offers a useful taxonomy of enterprise AI cost. It spreads across four layers: direct model and API spend; AI bundled into existing SaaS subscriptions; cloud and inference infrastructure; and agentic workloads. Each layer behaves differently, but they all share one quality: they are harder to forecast than traditional software costs. Direct model and API spend is usage-based, which means the same headcount can produce a 40%+ swing in monthly spend. AI bundled into SaaS is the least visible layer; your CRM, ERP, and dev tools now ship AI features at a higher price tier, often with no separate line item. Zylo found that 78% of IT leaders encountered unexpected charges from consumption-based and AI pricing. Cloud and inference infrastructure looks like classic cloud costs but with far more volatility. Agentic workloads are the newest and most unpredictable layer, often powered by autonomous systems whose token consumption cannot be reliably predicted before deployment. Dell Technologies senior vice president Varun Chhabra described what this feels like in practice: “You could kick off a workflow, come back the next morning and the workflow is done, but you don’t know how many tokens it consumed, whether it was efficient, or how much infrastructure it actually used.” “I’ve reached my token limit” has become a normal sentence in engineering organizations. That phrase, once reserved for consumer apps, now dictates whether a developer can finish their work. The infrastructure layer leaks almost as much. ClearML’s State of AI Infrastructure report found that 35% of enterprises rank increasing GPU utilization as their top priority, yet 44% say they manually assign workloads to GPUs or have no specific utilization strategy. In a world where GPU hours are the new budget currency, manual assignment is the equivalent of paying for a car with a coin roll. The architectural mismatch runs even deeper. JG Chirapurath, a former Azure VP and current chief marketing and solutions officer at SAP, described the bind: “Companies process only 20-30% of their available data because processing everything would blow their compute budgets by 5x to 10x. One Fortune 100 retailer I worked with had 15 years of customer interaction data but could only afford to process 30% of it. Their AI was essentially flying blind.” The data infrastructure built for overnight reports and batch jobs cannot absorb AI’s appetite for context.

The community has been saying this all along

The analysts are late to a party that developers have been watching for a while. A Hacker News thread that drew 594 upvotes put it bluntly: “Companies are burning through AI budgets faster than they can measure ROI, and the reason is foundational: they’re automating chaos instead of eliminating it.” The same discussion described enterprise AI acceleration as “structural chaos and technical debt,” driven less by productivity gains and more by FOMO. Reddit’s enterprise tech communities have developed a vocabulary for the failure modes: shadow AI spend, the demo effect, and token anxiety. One senior engineer on r/ExperiencedDevs wrote: “We shipped an AI feature that looked great in demo. In production, each user query cost $0.47. We had 50,000 users. Do the math. We killed it in three weeks.” GitHub issue trackers and discussions tell the same story from the tooling side. Teams complain they cannot predict inference costs for open-source model deployments, have no good way to monitor token consumption across development teams, and are being pushed into consumption-based pricing models that make budgeting nearly impossible. Then the real-world cases arrived. TechCrunch’s June 2026 reporting captured the moment when the meter caught up: Uber blew through its entire 2026 AI coding budget by April; Microsoft revoked its developers’ Claude Code licenses months after enabling them; one company reportedly racked up a $500 million Claude bill. “In April and May, I started hearing from companies saying, ‘We’ve exhausted our entire 2026 token budget.’” The Financial Times documented the same phenomenon inside Amazon. A project using Claude Sonnet to match authors to product pages failed after costing $1.8 million — 860% over budget — and the overrun was not noticed for five months. A finance-audit tool unexpectedly cost $541,000. A logistics-optimization feature cost $134,000 and was discovered only weeks later. Amazon engineers blamed deployment errors and the absence of cost limits. In a traditional system, the same kinds of mistakes would be cheap. In AI, mistakes compile into enormous token bills.

[SPONSORED]

COMFYUI WORKFLOW OPTIMIZATION

Reduce render times by 40% with our automated edge-silicon pipelines. Download Whitepaper.

The governance gap is the real gap

KPMG’s survey on AI cost transparency found that only 26% of companies say they have full visibility into their AI spend. 50% have partial visibility. 22% have none — some only discover the cost when the cloud bill arrives. KPMG’s global AI lead, Steve Chase, has described AI as a resource that enterprises never had to manage before and are still learning to control. CloudZero’s 2026 survey of 260 finance leaders landed on the same nerve: only 22% can link AI spending to business outcomes today, even though 87% say they need to by the end of the year. Harness estimates that for every $4 an enterprise spends on AI, about $1 is wasted. That is not a rounding error; it is the cost of ambiguity. The fixes are not exotic. FinOps Foundation guidance calls for clear AI ownership, cost tracking down to token or GPU level, and incremental funding that allows teams to fail fast and cheaply. The harder part is organizational. SpendHound found that a fifth of companies don’t know who owns the AI budget; that single fact explains the overrun rate better than any model architecture. Wipro’s global CIO Kenny Kesar has been candid about the lessons: without the right process orchestration, AI is a very expensive experiment. Plenty of use cases are “AI for AI’s sake,” and that is what breaks the economics. The fix is not to stop funding AI — it is to force every AI proposal to answer one question: what business process changes, and who is accountable for the change? If no one has the answer, the project should probably not exist.

The 2028 forecast is an accounting problem dressed up as a technology problem

Enterprise AI has not hit a ceiling in model quality. It has hit a ceiling in financial control. The 2028 Gartner forecast is not really about the future; overruns are already happening in the Amazon projects, the Uber budget, the Microsoft token caps, and the finance teams that cannot see their own bills. The only thing left to forecast is how much worse it gets. More capable models, better agents, and cheaper open-source weights will not solve the problem by themselves. If anything, they make the cost crisis easier to hide. The companies that emerge from this cycle with real AI advantage will not be the ones that bought the most tokens. They will be the ones that know, before the quarter ends, exactly what every token bought them. So ask the obvious question in your next budget review: who can tell you the actual cost of one production workflow, end to end, including retries, agents, and the SaaS apps that quietly bundle AI pricing in? If the answer is “no one,” the 2028 overrun forecast is not a prediction. It is your current plan.

Editorial Disclosure: This commercial analysis is compiled from global informational platforms and developer community discussions. Due to rapid technical cycles, readers are advised to independently verify volatile metrics. FUTUREMARSNEWS maintains structural objectivity and independent neutrality. more
This publication is intended solely for commercial, educational, and informational purposes. Articles may include news reporting, editorial opinions, technical analysis, software tutorials, deployment guidance, benchmark testing, hardware evaluations, workflow optimization strategies, pricing references, market intelligence, developer resources, and enterprise technology commentary. Product specifications, APIs, licensing models, cloud pricing, benchmark results, software capabilities, commercial terms, and hardware availability are subject to change without notice. Any performance figures or comparisons are based on publicly available information, vendor documentation, independent testing, or specific test environments and should not be interpreted as universally representative. Readers are encouraged to verify all technical and commercial information directly with official vendors before making engineering, purchasing, investment, or operational decisions. Unless explicitly labeled as sponsored content, advertising, affiliate content, or paid partnerships, editorial decisions remain independent. FUTUREMARSNEWS does not warrant the completeness, accuracy, or future availability of third-party products, services, software, or information referenced within this publication.