CompaniesInvestorsPeople
Home
Loading

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

Get in Touch

  • Contact

  • Request a Demo

  • Request Data Updates

  • Add a Company

Research

  • Companies

  • Investors

  • People

aVenture

  • Download App

  • Pricing

Download the aVenture Research beta for iOS and iPadOSDownload aVenture Research on the Mac App Store

Resources

  • Documentation

  • CLI

  • MCP

  • Feature Requests

  • Sitemap

Member

Backed by

© aVenture Investment Company, 2026. All rights reserved.

San Francisco, CA, USA

Privacy Policy · Terms of Service

aVenture Investment Company ("aVenture") is an independent research platform providing detailed analysis and data on startups, venture capital investments, and key industry individuals. It is not a registered investment adviser, broker-dealer, or investment advisor and does not provide investment advice or recommendations. The data provided by aVenture does not constitute recommendations or advice, whether by methodology, analysis, AI-generated content, or a statement written by a staff member of aVenture.

aVenture is not affiliated with any of the people, companies, organizations, government agencies, regulatory bodies, or investment funds we provide coverage for on this site unless explicitly stated otherwise. Users assume full responsibility for decisions made based on information obtained from this platform. Links to external websites do not imply endorsement or affiliation with aVenture. Any links that provide the ability to invest in a primary or secondary transaction in a company are for convenience only and do not constitute solicitations or offers to buy or sell an investment. Investors should exercise heightened precaution and due diligence when investing in private companies, especially those not independently audited.

While we strive to provide valuable insights with objectivity and professional diligence, we cannot guarantee the accuracy of the information provided on our platform. Before making any investment decisions, you should verify the accuracy of all pertinent details for your decision. To the fullest extent permitted by law, aVenture shall not be liable for any direct, indirect, incidental, consequential, or financial damages arising from use of this site, whether by consumers of its contents directly or by persons or organizations covered by our research, even if we are advised of the possibility. Our best-efforts processes and correction request forms do not create a warranty or duty of care.

Profiles on this platform may include content generated in part by large language models (LLMs, artificial intelligence) that aggregate publicly available sources (e.g., SEC EDGAR, public filings, press releases). Source attribution is provided where known; always verify statements and claims here against original sources before relying on any data. Content on our site may contain inaccuracies, omissions, or what are commonly called 'hallucinations' if generated in part or in full by AI / LLMs. The risk can also exist even when content is written by a human, as internal and third-party sources may also have inaccuracies for the same or different reasons. While we randomly audit a proportion of content, this is not exhaustive.

We recommend that an independent auditor be hired to verify the accuracy of the information before relying on it for any sensitive decisions. By accessing this platform, you agree not to rely solely on any information generated by AI, aggregated, or sourced or written otherwise on this site, for investment, financial, or other decisions. aVenture assumes no responsibility for inaccuracies, omissions, or hallucinations. You must independently verify all data from primary sources. Use of this platform constitutes your waiver of claims for reliance-based damages, including negligent misrepresentation. To report an error, request a correction, or dispute information about a company or individual, contact us via our request data updates form.

Loading
Loading
Home
News
Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

From SiliconAngle

By Zeus Kerravala

September 28, 2026

Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

Agentic AI is breaking the token meter, and enterprises need a plan for what comes next

Per-token pricing was the best thing to happen to enterprises looking to experiment with artificial intelligence, but it may be the worst thing for AI in production.

That’s the quandary at the center of a new Futurum report, “The Off Ramp From Per-Token Pricing,” sponsored by neocloud provider QumulusAI Inc. The report’s key finding is that agentic AI can drive token consumption per task 10 to 100 times higher than a simple inference call. Agentic AI being more expensive is likely no surprise, but it’s good to see the report quantify it.

This is a real problem I have heard from chief information officers and chief financial officers over and over with increasing frequency. The most successful AI projects end up costing the most, often with budget estimates way off. I spoke with one organization that budgeted $1 million for the year, and the initiative was so successful they spent it in three months.

Usage pricing punishes success

The appeal of per-token pricing is obvious. A developer can call an application programming interface and have a working prototype by the afternoon, without any capacity planning or procurement cycle. The problem is that the meter doesn’t distinguish between a pilot and a production system serving 20,000 employees, and agents are token machines. A chatbot answers a question. An agent plans, calls tools, checks its work, retries, hands off and summarizes, generating tokens at each step.

Futurum forecasts that agent and reasoning inference will grow by 219% this year, and total inference spending will rise from $120 billion in 2025 to $885 billion by 2030. Put those numbers next to a pricing model that scales linearly with consumption, and the result is a budget line that grows faster than the value it creates.

Mazda Marvasti, co-founder and chief executive of Amberd.ai, described the pattern in the report. “When they start deploying it throughout the organization, the cost starts skyrocketing because it’s a useful tool that somebody built, but it’s now priced on a variable basis,” he said. “It starts getting the attention of the CFO and the CIO in terms of how much I’m exactly spending to run this tool, and whether it’s worth it.”

That captures the overlooked part of this story. The risk isn’t just a high bill; it’s that unpredictable bills kill useful projects. Marvasti noted that some customers abandoned internally built automation tools because they couldn’t forecast or justify the costs. That’s a governance failure disguised as a pricing problem, and it will slow AI adoption more than any model limitation.

Enterprises have already voted with their capacity

A more interesting data point is that the market has quietly moved beyond the “everything in the public cloud” assumption. According to Futurum’s survey of 824 AI decision-makers, reserved and owned infrastructure account for 66% of AI compute consumption, compared with 19% for on-demand cloud. Some 59% of respondents primarily run AI workloads outside hyperscaler public clouds, in their own data centers, colocation facilities, or with bare-metal providers.

I’d caution against interpreting this as enterprises moving away from the hyperscalers. Much of that owned capacity reflects GPU purchases made when on-demand capacity simply wasn’t available. But it shows enterprises are comfortable making capacity commitments for AI, and the question is no longer whether to commit, but which workloads justify such commitments.

This mirrors the adoption cycle information technology went through with the cloud. Start on demand, discover that steady-state workloads are cheaper on reserved capacity, and end up hybrid. AI is compressing that curve from years to quarters, and agents are the accelerant.

The real economics are about utilization, not price

One of the more interesting sections of the report is Amberd.ai’s deployment on QumulusAI bare metal. The company partitions an eight-GPU Nvidia H200 server into four virtual environments, each with two GPUs, and tiers customers across them based on latency tolerance.

“With one 8x H200 server, two customers pay for the entire server, and I can probably have about 30 to 35 customers running on that one server,” Marvasti said. “After the second customer, the server is free to me, and any customer that comes after that is profit.” Though those data points are compelling, the lesson isn’t that bare metal is cheap. It’s that Amberd.ai built a custom virtualization layer and tiered pricing to drive utilization. Reserved infrastructure turns a variable cost into a fixed one, and fixed costs only pay off when kept busy. An idle reserved GPU is the most expensive GPU there is.

Futurum’s guidance supports this, recommending reserved bare metal for sustained workloads with predictable utilization above roughly 60%. The report also acknowledges that these environments “require more custom engineering, limiting the operating margin gains for teams without the hardware expertise.”

That caveat deserves more attention than it receives. Most enterprises lack deep bench strength in serving engines, batching, quantization and key-value cache management. Without it, the offramp can lead to a different kind of cost overrun.

The model question the report doesn’t answer

The offramp works well for open-weight models a company can deploy on infrastructure it controls, which is why Amberd.ai built on private, open-source large language models. But many enterprises have standardized on frontier models available only through their developers’ application programming interfaces or hyperscaler marketplaces. For those workloads, no bare-metal alternative exists, so the pricing lever rests with the model provider.

That makes the reserved-versus-per-token decision a model-strategy decision first. Open-weight models are gaining momentum and are increasingly good enough for classification, extraction, summarization and many agent subtasks. The companies with the most leverage will route work to the cheapest model that does the job well, then run it on the cheapest infrastructure that runs reliably.

In the report, Brennen Smith, chief technology officer of Runpod, explained the importance for finance teams. “If you have a predictable workload, such as a well-defined business operation, a fixed lease contract is the way to go,” he said. “However, if it’s experimentation, scaling, or variable velocity, that’s when you need to go on demand. When talking to CFOs, I recommend budgeting for both.”

It’s the right answer, but a harder organizational change than it sounds, because it requires finance, infrastructure, and AI teams to share a view of workload behavior that most companies lack today.

What this means for buyers

Per-token pricing isn’t going away, and it shouldn’t. But treating it as the default for production AI is a mistake that agentic workloads will quickly reveal. My advice for IT and finance leaders:

  • Measure cost per task, not cost per token. Agents change the unit of work. As agent complexity grows, track the cost to resolve a ticket or process a claim end to end.
  • Set a graduation trigger. Define the utilization and volume thresholds that determine when a workload transitions from APIs to on-demand GPUs and then to reserved capacity. Futurum’s 60% baseline is a reasonable starting point.
  • Be honest about operational skills. Reserved infrastructure only saves money if it stays busy. If that expertise isn’t in-house, factor in managed services before committing.
  • Test open-weight models now. Every workload that can run on an open model can move off the meter.
  • Negotiate for hardware cycles. Contracts should address upgrade paths, renewals, and portability so today’s commitment doesn’t become tomorrow’s stranded capacity.

The companies that win with agentic AI won’t necessarily be the ones with the best models. They’ll be the ones that figured out how to afford to run them at large scale.

Zeus Kerravala is a principal analyst at ZK Research, a division of Kerravala Consulting. He wrote this article for SiliconANGLE.

View original article on siliconangle.com

Most Recent

1Password ties AI agent access to individual tasks

1Password ties AI agent access to task-level checks, just-in-time permissions and credentials kept outside the underlying model.

Sep 28, 2026

Meta hires MongoDB CEO CJ Desai to lead new enterprise AI business

Meta Platforms Inc. is launching a new business unit that will provide artificial intelligence services to enterprises. The Meta Enterprise Platform, as it’s called, will be led by longtime technology executive CJ Desai. The company stated in a launch announcement today that he will hold the title o

Sep 28, 2026

Qiagen grounds drug discovery agents in curated knowledge

Qiagen explains why AI agents in drug discovery need traceable, human-curated knowledge to avoid hallucinated outputs, per theCUBE interview.

Sep 28, 2026

Seattle-based sales tech company Outreach to move HQ to Adobe’s campus in Fremont neighborhood

Outreach confirmed plans to move its headquarters from Interbay to Seattle’s Fremont neighborhood early next year, joining a campus that Adobe has anchored since the late 1990s. The sales technology company, which was valued at more than $4.4 billion in a 2021 funding round, has grown to about 800 e

Sep 28, 2026

Similar Posts

1Password ties AI agent access to individual tasks

1Password ties AI agent access to task-level checks, just-in-time permissions and credentials kept outside the underlying model.

Sep 28, 2026

Meta hires MongoDB CEO CJ Desai to lead new enterprise AI business

Meta Platforms Inc. is launching a new business unit that will provide artificial intelligence services to enterprises. The Meta Enterprise Platform, as it’s called, will be led by longtime technology executive CJ Desai. The company stated in a launch announcement today that he will hold the title o

Sep 28, 2026

ServiceNow calls for a measured response to rogue AI agents

ServiceNow frames agent containment as a graduated response guided by risk, permissions and the business processes an AI agent supports.

Sep 28, 2026

Qiagen grounds drug discovery agents in curated knowledge

Qiagen explains why AI agents in drug discovery need traceable, human-curated knowledge to avoid hallucinated outputs, per theCUBE interview.

Sep 28, 2026