Data has become the fuel for artificial intelligence, and the largest untapped reserve to power AI is unstructured data.
Industry estimates cited by theCUBE Research indicate that more than 80% of enterprise data is unstructured, yet this critical pool remains largely inaccessible, unmanaged or underutilized. This data is generally composed of documents, videos, logs, emails, images, audio recordings, and other information that could go a long way toward improving AI-generated results.
The challenge lies in discovering, governing and ultimately moving this data for AI consumption, a task that several key companies in the storage arena are addressing today.
“This idea of managing unstructured data for AI is much more challenging than managing unstructured data in the past where it was kind of a little bit more about just archive and maybe some backup and recovery types of initiatives,” said Molly Presley (pictured, top right), chief marketing officer of Hammerspace Inc. “The idea of unifying and then really efficiently automating the movement of it are big pieces of this unstructured data management for AI.”
Presley spoke with Rob Strechay for the Supermicro Open Storage Summit interview series, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. She was joined by Sherry Lin (bottom, left), senior product manager of SDS solution at Super Micro Computer Inc.; Peter Sjoberg (top, left), vice president of worldwide solution architects at Cloudian Inc.; and Mohamad El-Batal (bottom, right), chief systems technologist, Office of the CTO, at Seagate Technology LLC. They discussed how organizations are managing unstructured data at scale and why effective data preparation has become one of the most important foundations for successful AI deployments. (* Disclosure below.)
Unstructured data management solutions
Supermicro has approached the management of unstructured data at scale by designing high-density hardware platforms that support software-defined storage solutions and ecosystem partner integrations. The company has also introduced Context Memory storage servers designed to support the offloading and sharing of large language model key-value caches across AI inference infrastructure.
“We offer systems for KV cache tiers that require high throughput and ultra low latency for most inferencing workloads,” Lin explained. “We co-engineered and validated our systems with our ecosystem partners like Hammerspace, Cloudian and Seagate. For example, Hammerspace will provide the global unified namespace and data orchestrations among the data from each tier. And Cloudian will [serve] the S3 data lake for object storage for AI applications. And that data, probably in the end of the life cycle, will sit in Seagate’s hard drive.”
Cloudian’s solution involves enterprise-grade, S3-native object storage engineered to manage and scale massive volumes of unstructured data. A central part of the company’s approach is to manage this information within a robust governance framework, according to Sjoberg.
“We see a key goal to put that unstructured data under management so that it is protected, it is safe and secure,” Sjoberg told theCUBE. “That’s exactly what we’re doing now as this data has moved into the AI era. That data will need to move for different purposes, and you will need to make sure it’s always under your control, unstructured or otherwise, so that you can bring it to your AI needs.”
Understanding infrastructure needs
The process of managing unstructured data for AI has brought greater focus to the value inherent in preparing it for future training, as well as for compliance audits, debugging and understanding AI outputs. This process is also an important step in determining the kind of infrastructure organizations will need, as Seagate’s El-Batal pointed out.
“The biggest challenge is going to be identifying the value of your data or being aware of where you are you going with that data, and what structures you want to build in an AI model to leverage the key aspect of productization that you want to go after,” El-Batal said. “That basically tells you what balance to pick in your infrastructure. The idea is to value your data and don’t go cheap on your infrastructure, while at the same time being efficient at it.”
Being efficient also means orchestrating infrastructure to generate results from unstructured data as rapidly as possible. For Hammerspace, this includes work around a Model Context Protocol layer designed to give AI systems greater understanding of and access to distributed data for model-driven pipelines and agents
“This is about creating your data estate, curating it, and then when you run the GPUs, going as fast as you can,” Presley said. “It’s just a very different workflow. On the Hammerspace side, we are really focused on that, the new MCP layer where you can get understanding of the data that’s sitting out there because it’s not well described. That’s the very nature of unstructured data.”
Stay tuned for the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the Supermicro Open Storage Summit interview series.
(* Disclosure: TheCUBE is a paid media partner for the Supermicro Open Storage Summit interview series. Neither Supermicro, the sponsor of theCUBE’s event coverage, nor other sponsors have editorial control over content on theCUBE or SiliconANGLE.)





