IBM and CoreWeave Co-Design Controls for Agent Workloads
For years, the main job of an AI research cluster was training models. That job is getting broader. Once teams use reinforcement learning (RL) and start testing agents, the same infrastructure has to run code that models produce, call tools, read and write storage and connect to other services. Pure training never raised the question of where that code should run and what it should be allowed to reach. Agent work does.
IBM Research is working through that question with CoreWeave, the GPU cloud provider. Brian Belgodere, a senior technical staff member at IBM, described the collaboration in an interview with theCUBE Research's Dave Vellante and John Furrier at the Fully Connected event. The interview was broadcast on theCUBE, the livestreaming studio of SiliconANGLE Media. theCUBE is a paid media partner for the event, and CoreWeave sponsored its coverage. SiliconANGLE says sponsors have no editorial control over its content.
Why reinforcement learning changes the cluster
Belgodere explained that RL adds a task-execution stage to model development. "The RL process is [that] you are in the middle of training a model," he said. "At some point, you take that checkpoint and then actually load it into inference, ask it to do something and you're measuring. That's your testing phase."
In practice, the training loop now contains an inference step and an execution step. A checkpoint is loaded, given a task and scored on the result. For agentic research, that task can mean running code. That's a different risk profile from a GPU doing matrix math, and it has to be planned for at the infrastructure level.
From a self-built H100 cluster to CoreWeave
IBM's infrastructure story starts with Granite, its family of AI models. Developing them took a lot of computing power, and at first IBM built that capacity itself. "We went out and built a large H100 cluster, and we did it ourselves: got the space, soup to nuts. It was a huge task," Belgodere said.
The next hardware generation brought cooling and power demands that, according to Belgodere, helped push IBM toward working with CoreWeave instead of repeating the do-it-yourself approach.
Joint engineering on identity
The relationship has since moved past renting capacity. IBM gave CoreWeave requirements for extending its internal identity systems into the CoreWeave environment. The two companies then refined the implementation over several iterations.
Belgodere said the feedback runs in both directions. "Oftentimes they will come to us and say, 'Hey, we're thinking of this, can we get your feedback?'" he said. "And we'll happily go through it."
This matters for enterprise buyers. Large organizations rarely want a separate identity system for each cloud provider. Pushing existing corporate identity into the provider's platform keeps access control in one place.
Choosing where agent code runs
Most of IBM Research's cluster is single-tenant. That means the hardware is not shared with other customers. IBM has also deployed its own storage inside CoreWeave, and it can tap additional capacity as long as it stays within cost and security parameters.
The collaboration also covers CoreWeave Sandboxes. These support isolated execution either on dedicated infrastructure or through a managed serverless runtime. Researchers can decide where agent code runs and which resources it can touch.
Belgodere warned that teams often get these early choices wrong. "There are a lot of misses, and people tend to underestimate the cost to change some of these decisions," he said. "If you decide to make a poor architecture decision early on, the cost is either going to be [that] you accidentally bought way too much networking infrastructure, or you have to go buy and refit everything."
Measuring what security costs
IBM does not treat security controls as free. It measures their performance impact against benchmark results and uses those numbers in its discussions with security teams about tradeoffs. Workload isolation is one part of that security architecture, next to the enterprise identity integration.
Belgodere framed the larger challenge as one of provenance. "This is a supply chain problem, top to bottom," he said. "It's not just the hardware, it's the firmware, kernel levels, code, your data provenance. Then you get into the whole world of agents, your images. It is an absolute provenance problem."
The Bigger Picture
The interview shows that agent safety is moving down the stack. A lot of public discussion focuses on model behavior: what an agent says, what it refuses, how it reacts to prompt injection. IBM's approach suggests that for teams running agents at scale, the boundary that matters most may be physical and architectural. That means which machine the code runs on, which identity it holds and which storage it can see.
This fits a wider pattern. Recent incidents, such as agents leaking company screenshots on GitHub, show what can happen when agents get more access than they need. CoreWeave is also building out security partnerships elsewhere, including its work with CrowdStrike. IBM has made similar moves on the tooling side by offering a self-hosted version of IBM Bob for air-gapped environments.
Two things are worth watching. First, whether sandboxed execution becomes a standard feature that GPU clouds compete on rather than an add-on. Second, whether IBM's habit of benchmarking the performance cost of each control spreads. If more teams publish such numbers, the security-versus-speed debate could rest on data instead of assumptions.
