AI Business

Agentic AI's Token Appetite Forces Enterprises to Rethink Pricing Models

A new report shows that agentic AI systems consume 10 to 100 times more tokens per task than simple inference, making per-token pricing unsustainable for production workloads and pushing enterprises toward reserved capacity and open-weight models.

·6 min read
Agentic AI is breaking the token meter, and enterprises need a plan for what comes next
Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

Per-token pricing democratized AI experimentation for enterprises, but it may prove catastrophic for systems running at scale. This tension sits at the heart of a Futurum report titled "The Off Ramp From Per-Token Pricing," underwritten by neocloud provider QumulusAI Inc. The study's central finding reveals that agentic AI workloads can generate 10 to 100 times more token consumption than straightforward inference operations. While the higher expense of agentic systems may seem obvious, quantifying the magnitude matters.

Chief information officers and chief financial officers have raised this concern with growing urgency. Successful AI initiatives frequently exceed budget projections dramatically. One organization allocated $1 million for the year but exhausted the entire budget in just three months due to project success.

Usage pricing punishes success

Per-token pricing's attraction is straightforward: developers can build working prototypes through API calls without capacity planning or procurement delays. The flaw emerges when the same meter applies equally to pilots and production systems serving tens of thousands of users. Agents function as token generators by design. While a chatbot responds to queries, an agent performs planning, invokes tools, validates results, retries operations, delegates tasks, and produces summaries—each step consuming tokens.

Futurum projects that agent and reasoning inference will expand by 219% this year, with total inference spending climbing from $120 billion in 2025 to $885 billion by 2030. When paired with linear pricing models, consumption-based costs grow faster than organizational value.

Mazda Marvasti, co-founder and chief executive of Amberd.ai, articulated this dynamic in the report: "When they start deploying it throughout the organization, the cost starts skyrocketing because it's a useful tool that somebody built, but it's now priced on a variable basis. It starts getting the attention of the CFO and the CIO in terms of how much I'm exactly spending to run this tool, and whether it's worth it."

The underlying risk extends beyond sticker shock. Unpredictable bills terminate promising projects. Marvasti noted that some customers discontinued internally developed automation tools because cost forecasting became impossible. This represents a governance problem masquerading as a pricing issue, and it will constrain AI adoption more effectively than any technical limitation.

Enterprises have already voted with their capacity

The market has quietly shifted away from the "all public cloud" paradigm. According to Futurum's survey of 824 AI decision-makers, reserved and owned infrastructure accounts for 66% of AI compute consumption, versus 19% for on-demand cloud services. Fifty-nine percent of respondents primarily execute AI workloads outside hyperscaler public clouds, operating instead in private data centers, colocation facilities, or through bare-metal providers.

This shouldn't be misread as enterprises abandoning hyperscalers entirely. Much of this owned capacity reflects GPU acquisitions made when on-demand options were unavailable. Nevertheless, it demonstrates that enterprises accept capacity commitments for AI, and the conversation has shifted from whether to commit to which workloads warrant such commitments.

This mirrors the adoption trajectory information technology experienced with cloud computing. Organizations begin on-demand, discover that steady-state workloads cost less on reserved capacity, and eventually operate hybrid environments. AI is compressing this cycle from years into quarters, with agents serving as the catalyst.

The real economics are about utilization, not price

The report highlights Amberd.ai's deployment on QumulusAI bare metal as particularly instructive. The company divides an eight-GPU Nvidia H200 server into four virtual environments containing two GPUs each, then distributes customers across these partitions based on latency requirements.

"With one 8x H200 server, two customers pay for the entire server, and I can probably have about 30 to 35 customers running on that one server," Marvasti explained. "After the second customer, the server is free to me, and any customer that comes after that is profit." While these figures appear compelling, the takeaway isn't that bare metal is inexpensive. Rather, Amberd.ai engineered a custom virtualization layer and implemented tiered pricing to maximize utilization. Reserved infrastructure converts variable costs into fixed ones, and fixed costs only generate returns when the hardware remains productive. An underutilized reserved GPU becomes the costliest option available.

Futurum's recommendations align with this principle, suggesting reserved bare metal for sustained workloads demonstrating predictable utilization above roughly 60%. The report also notes that such environments "require more custom engineering, limiting the operating margin gains for teams without the hardware expertise."

This caveat warrants greater emphasis. Most enterprises lack substantial expertise in serving engines, batching, quantization, and key-value cache management. Without this knowledge, the alternative path can lead to different cost overruns.

The model question the report doesn't answer

The alternative works effectively for open-weight models that organizations can deploy on their own infrastructure, which explains why Amberd.ai selected private, open-source large language models. However, many enterprises have committed to frontier models accessible only through their creators' APIs or hyperscaler marketplaces. For these workloads, bare-metal alternatives don't exist, leaving pricing control with the model provider.

This makes the reserved-versus-per-token choice fundamentally a model-strategy decision. Open-weight models are gathering momentum and have reached sufficient quality for classification, extraction, summarization, and numerous agent subtasks. Organizations with the greatest flexibility will direct work to the most cost-effective model that performs adequately, then execute it on the most economical infrastructure that operates reliably.

Brennen Smith, chief technology officer of Runpod, addressed this for finance teams in the report: "If you have a predictable workload, such as a well-defined business operation, a fixed lease contract is the way to go. However, if it's experimentation, scaling, or variable velocity, that's when you need to go on demand. When talking to CFOs, I recommend budgeting for both." This represents sound guidance, though executing it requires organizational change more substantial than it appears, demanding that finance, infrastructure, and AI teams align on workload characteristics most organizations haven't yet characterized.

What this means for buyers

Per-token pricing will persist, and rightfully so. However, treating it as the standard for production AI systems represents a miscalculation that agentic workloads will quickly expose. Guidance for IT and finance leaders includes:

  • Measure cost per task, not cost per token. Agents redefine the unit of work. As agent sophistication increases, monitor the cost to resolve a ticket or process a claim from beginning to end.
  • Set a graduation trigger. Establish the utilization and volume thresholds that signal when a workload moves from APIs to on-demand GPUs and subsequently to reserved capacity. Futurum's 60% baseline offers a reasonable reference point.
  • Be honest about operational skills. Reserved infrastructure only delivers savings if kept busy. If that expertise doesn't exist internally, include managed services costs before making commitments.
  • Test open-weight models now. Any workload capable of running on an open model can exit the per-token meter.
  • Negotiate for hardware cycles. Contracts should specify upgrade paths, renewal terms, and portability provisions so current commitments don't become future stranded assets.

Enterprises that succeed with agentic AI won't necessarily possess the superior models. Rather, they'll be those that mastered the economics of operating them at enterprise scale.