CompaniesInvestorsPeople
Home
Loading

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

aVenture is in Beta: research coverage is expanding as we build, so please independently verify key details before making investment decisions.

Get in Touch

  • Contact

  • Request a Demo

  • Request Data Updates

  • Add a Company

Research

  • Companies

  • Investors

  • People

aVenture

  • Download App

  • Pricing

Download the aVenture Research beta for iOS and iPadOSDownload aVenture Research on the Mac App Store

Resources

  • Documentation

  • Use Cases

  • CLI

  • MCP

  • Feature Requests

  • Sitemap

Member

Backed by

Ask AI about aVenture

© aVenture Investment Company, 2026. All rights reserved.

San Francisco, CA, USA

Privacy · Terms of Service

aVenture Investment Company ("aVenture") is an independent research platform providing detailed analysis and data on startups, venture capital investments, and key industry individuals. It is not a registered investment adviser, broker-dealer, or investment advisor and does not provide investment advice or recommendations. The data provided by aVenture does not constitute recommendations or advice, whether by methodology, analysis, AI-generated content, or a statement written by a staff member of aVenture.

aVenture is not affiliated with any of the people, companies, organizations, government agencies, regulatory bodies, or investment funds we provide coverage for on this site unless explicitly stated otherwise. Users assume full responsibility for decisions made based on information obtained from this platform. Links to external websites do not imply endorsement or affiliation with aVenture. Any links that provide the ability to invest in a primary or secondary transaction in a company are for convenience only and do not constitute solicitations or offers to buy or sell an investment. Investors should exercise heightened precaution and due diligence when investing in private companies, especially those not independently audited.

While we strive to provide valuable insights with objectivity and professional diligence, we cannot guarantee the accuracy of the information provided on our platform. Before making any investment decisions, you should verify the accuracy of all pertinent details for your decision. To the fullest extent permitted by law, aVenture shall not be liable for any direct, indirect, incidental, consequential, or financial damages arising from use of this site, whether by consumers of its contents directly or by persons or organizations covered by our research, even if we are advised of the possibility. Our best-efforts processes and correction request forms do not create a warranty or duty of care.

Profiles on this platform may include content generated in part by large language models (LLMs, artificial intelligence) that aggregate publicly available sources (e.g., SEC EDGAR, public filings, press releases). Source attribution is provided where known; always verify statements and claims here against original sources before relying on any data. Content on our site may contain inaccuracies, omissions, or what are commonly called 'hallucinations' if generated in part or in full by AI / LLMs. The risk can also exist even when content is written by a human, as internal and third-party sources may also have inaccuracies for the same or different reasons. While we randomly audit a proportion of content, this is not exhaustive.

We recommend that an independent auditor be hired to verify the accuracy of the information before relying on it for any sensitive decisions. By accessing this platform, you agree not to rely solely on any information generated by AI, aggregated, or sourced or written otherwise on this site, for investment, financial, or other decisions. aVenture assumes no responsibility for inaccuracies, omissions, or hallucinations. You must independently verify all data from primary sources. Use of this platform constitutes your waiver of claims for reliance-based damages, including negligent misrepresentation. To report an error, request a correction, or dispute information about a company or individual, contact us via our request data updates form.

Loading
Loading
Home
News
AI inference gets a new tier as context windows grow

From SiliconAngle

By Victoria Gayton

August 25, 2026

AI inference gets a new tier as context windows grow

AI inference gets a new tier as context windows grow

AI storage infrastructure is becoming a more consequential planning issue as organizations move from model training toward agentic AI. As agents reason, act and reassess, they build longer contexts and generate more data that they must access quickly during inference.

Agentic AI is also changing the shape of the data problem. Interactions are growing longer and producing more information. At the same time, the data’s size, importance and movement through the infrastructure can all affect how quickly an application responds, according to Scott Shadley (pictured, left), director of technology planning at Solidigm Inc.

“One of the beautiful things that’s happened in this agentic AI, or even just the AI era, is [that] people are starting to pay attention to storage,” he said. “What’s unique about this particular era is it’s no longer one- or two-dimensional. We have data magnitude and growth in size, importance and all of the other volumetric aspects of that. But at the end of the day, it comes down to that bit of data and how fast that bit of data moves from point A to point B.”

Shadley, along with Anat Heilper (center), director of AI architecture at Vast Data Inc., and Ben Lee (right), director of solution management at Super Micro Computer Inc., spoke with theCUBE Research’s Rob Strechay during the Supermicro Open Storage Summit interview series. They discussed how the companies’ respective technologies fit together as expanding context windows and KV caches create new storage and memory demands for agentic AI.

AI storage infrastructure brings KV cache closer to compute

As context windows expand, graphics processing unit memory alone can’t hold everything an agentic workload needs during inference. AI storage infrastructure must therefore provide additional tiers that balance proximity, capacity and speed, with each layer handling a different part of the data load, according to Shadley.

“As you think through that architecture — and you need to put that context somewhere — that context can start living in what used to be a no-no zone,” he said. “So [solid-state drives] have found a new home. One of the unique things about this 3.5 tier that we’re creating is that it could not exist until we had things like [Non-Volatile Memory Express] SSDs.”

That middle tier is one part of a larger architecture. Solidigm supplies SSDs that provide fast access to cached data, Supermicro integrates them into rack-scale systems and Vast Data’s AI Operating System provides network storage and the data services needed to use, manage and protect that data alongside broader AI workloads, Heilper noted.

“When we talk about [key-value] cache, which is a very significant optimization that can be done in AI inferencing … in essence, it’s the ability to replace compute with storage,” she said. “This is very significant because we all know the GPU is very, very expensive. When you have very high KV cache hit rates, we both save on compute and reduce the latency significantly.”

AI infrastructure needs room to evolve

Supermicro’s Context Memory eXtension, or CMX, proposal targets organizations with large AI clusters and substantial data demands; other deployments may require different combinations of memory, local SSDs and network storage. The company’s broader value lies in composing those building blocks around each customer’s workload, rather than treating one architecture as a universal answer, according to Lee.

“We believe solving the problem will take the whole rack because you cannot just buy more GPUs with more [high-bandwidth memory] … it’s very expensive,” he said. “All the KV cache will naturally overflow from the GPU, HBM, to the system memory, to the local SSD and to the network storage. But there’s a new industrial definition to try and fill the gap, and they call it G3.5, which is the CMX solution. This is a very AI-native KV cache tier that can fulfill the demand.”

Vast Data has been testing KV cache offload with partners across different software environments. Those experiments aim to show how the architecture behaves when cached context is integrated into production inference workloads, according to Heilper.

“With Nvidia Dynamo, we’ve shown that we managed to get 20 times faster time-to-first-token, which means that the latency that you perceive as a user is significantly faster,” she said. “Also, [we’ve] seen 90% savings in GPU time.”

Those results depend on an AI storage infrastructure that can match storage performance and capacity to the workload. Solidigm’s D7-PS1010 performance-oriented SSD and D5-P5336 capacity drive address different points in the hierarchy, according to Shadley. Supermicro integrates those components into systems that can be configured around customer requirements.

“This is not the only definition of a hierarchy stack,” Shadley said. “It’s the current primary that everybody leverages as the gold standard, but it’s continuing to evolve, and it’s unique. Being very proactive with your customer or your supplier to better understand what they know about what you need is no longer transactional. Those value-level transaction conversations are now what are going to drive the future of us deploying these types of architectures.”

Stay tuned for the complete video, part of SiliconANGLE and theCUBE’s coverage of the Supermicro Open Storage Summit interview series.

(* Disclosure: TheCUBE is a paid media partner for the Supermicro Open Storage Summit interview series. Neither Supermicro, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)

Photo: SiliconANGLE

View original article on siliconangle.com

Most Recent

What to expect during the AI Data Pipeline Forum: Join theCUBE Oct. 13

AI infrastructure bottlenecks take focus at the AI Data Pipeline Forum Oct. 13. Get a preview of theCUBE's coverage of storage, networking and power.

Oct 10, 2026

Billionaire Jeff Bezos: ‘Success is not villainy’

When Jeff Bezos sat down on Blue Origin’s Cape Canaveral factory floor for an interview with Fox News’ Bret Baier earlier this week, he laid out an unapologetic defense of free enterprise and wealth creation. “Success is not villainy,” Bezos said. The Amazon founder’s point was simple: building a ma

Oct 9, 2026

Oxide Computer raises $445M to step up data center rack production

Data center hardware startup Oxide Computer Inc. today announced that it has raised $445 million in funding. Returning backer Eclipse Capital led the Series D round. It was joined by AMD Ventures, Riot Ventures, Jane Street, Atreides Management and USIT, which led Oxide’s previous $200 million raise

Oct 9, 2026

‘Star Wars’ is crossing over with ‘D&D’ in two upcoming books from Wizards of the Coast

Wizards of the Coast plans to revisit the Star Wars universe as its next big cross-company collaboration for Dungeons & Dragons. On the first day of New York Comic Con 2026, one of the reveals at its big Disney/Lucasfilm panel was the upcoming debut of two new sourcebooks for D&D based on the Star W

Oct 9, 2026

Similar Posts

The AI storage stack gets an inference-era rethink

Artificial intelligence is changing what storage and data management platforms look like. In collaboration with Super Micro Computer Inc. and Solidigm, DataDirect Networks Inc. has introduced DDN Enterprise AI HyperPOD, built on Nvidia Corp.’s AI Data Platform. The goal is to simplify the storage, s

Aug 27, 2026

AI data growth drives demand for hybrid storage architectures

The growth of AI workloads is changing how organizations approach data management, and hybrid storage is emerging as a key resource. While the cloud originally promised simplicity, the reality is that infrastructure has become more complex. Most organizations now operate across a combination of on-p

Sep 1, 2026

Cloudera brings Mistral AI’s frontier models into its secure hybrid data environments

Big-data company Cloudera Inc. is pushing to become the go-to partner for enterprise-grade sovereign artificial intelligence workloads after announcing a massive, nine-figure strategic partnership with the French AI model maker Mistral AI SAS. The partnership is aimed at bringing secure, sovereign A

Sep 9, 2026

AI storage platform Vast Data aimed for $25B valuation in new round, sources say

The AI-friendly data storage startup is raising capital at a giant leap in valuation.

Jun 10, 2025