Agent workload isolation is becoming a practical infrastructure challenge as research teams move beyond training models to running code and testing agents. Systems built for intensive computation must now accommodate workloads that interact with tools, storage and other services.
IBM Research’s infrastructure must support increasingly varied workloads as its model development practices evolve. Reinforcement learning adds a task-execution stage that changes what those systems must accommodate, according to Brian Belgodere (pictured), senior technical staff member at International Business Machines Corp.
“The RL process is [that] you are in the middle of training a model,” he said. “At some point, you take that checkpoint and then actually load it into inference, ask it to do something and you’re measuring. That’s your testing phase.”
Belgodere spoke with theCUBE Research’s Dave Vellante and John Furrier at the Fully Connected event, during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. They discussed evolving research workloads, security-performance tradeoffs, shared infrastructure engineering and agent workload isolation. (* Disclosure below.)
From model training to executing agent code
IBM Research’s work on its Granite family of models required substantial computing infrastructure. The cooling and power demands of a subsequent hardware generation helped drive IBM’s decision to work with CoreWeave Inc., according to Belgodere.
“That was what gave me, ‘Hey, go build out our infrastructure,’” he said. “We went out and built a large H100 cluster, and we did it ourselves: got the space, soup to nuts. It was a huge task.”
The relationship has expanded into joint engineering around identity management and workload controls. IBM supplied requirements for extending its internal identity systems into CoreWeave, then refined the implementation through several iterations, according to Belgodere.
“Oftentimes they will come to us and say, ‘Hey, we’re thinking of this, can we get your feedback?’” he said. “And we’ll happily go through it.”
Agent workload isolation becomes a design requirement
Much of IBM Research’s cluster is single-tenant, with its own storage deployed inside CoreWeave and additional capacity available within cost and security parameters, according to Belgodere. IBM’s collaboration with CoreWeave also includes CoreWeave Sandboxes, which support isolated execution on dedicated infrastructure or through a managed serverless runtime. These options let researchers choose where agent code runs and what resources it can access.
“There are a lot of misses, and people tend to underestimate the cost to change some of these decisions,” Belgodere said. “If you decide to make a poor architecture decision early on, the cost is either going to be [that] you accidentally bought way too much networking infrastructure, or you have to go buy and refit everything.”
IBM measures the performance impact of security controls against benchmark results, using those findings to discuss tradeoffs with security teams. Workload isolation sits alongside enterprise identity integration as part of that broader security architecture.
“This is a supply chain problem, top to bottom,” he said. “It’s not just the hardware, it’s the firmware, kernel levels, code, your data provenance. Then you get into the whole world of agents, your images. It is an absolute provenance problem.”
Here’s the complete video interview, part of SiliconANGLE’s and theCUBE’s coverage of the Fully Connected event:
(* Disclosure: TheCUBE is a paid media partner for the Fully Connected event. Neither CoreWeave, the sponsor of theCUBE’s coverage, nor other sponsors have editorial control over the content on theCUBE or SiliconANGLE.)
