CoreWeave Forge Bundles AI Inference and Post-Training Tools
CoreWeave Inc. wants to be more than a place to rent GPUs. At the Fully Connected event, the specialized cloud provider launched CoreWeave Forge, a platform that brings model serving, observability, post-training and evaluation together in one place. It also previewed a new capability called CoreWeave RL Rollouts, aimed at teams training agentic models with reinforcement learning.
The underlying argument is simple. Training built the first wave of GPU clouds. Serving models faster and at lower cost may decide who profits from the next one.
Inference is taking a bigger share of the work
Urvashi Chowdhary, vice president of product and AI services at CoreWeave, described the company's strategy in an interview with theCUBE Research's Dave Vellante and John Furrier. theCUBE is SiliconANGLE Media's livestreaming studio. In her view, AI developers want to solve a problem quickly, with strong performance and costs that hold up at scale. CoreWeave's response has been to add managed services for training, post-training and inference on top of its infrastructure.
The demand data points the same way. A survey of CoreWeave customers and prospects by theCUBE Research found that one healthcare customer's share of inference workloads grew from about 10% in its first year to 40% in its second. That share is expected to reach roughly 50% within 12 months.
For a provider that grew up on training contracts, that is a meaningful shift. It moves the competition away from raw GPU capacity and toward storage, networking and software.
Tuning every layer above the hardware
According to Chowdhary, CoreWeave is optimizing each layer of the inference stack that sits above the chips. That includes:
- the vLLM serving engine
- quantized models
- custom speculative decoders
She stressed that the company has deliberately leaned on open source tools and contributes back to open systems. The stated goal is to give customers flexibility while CoreWeave builds its own services on top of one another.
Where reinforcement learning hurts
Reinforcement learning creates a specific problem. When customers train agentic models with rewards and verifiers, the model keeps producing new checkpoints. Each new version has to be pushed into the inference setup during rollouts, and Chowdhary said inference becomes the bottleneck at that stage.
CoreWeave RL Rollouts is meant to address this. The preview capability is built on Nvidia Corp.'s Dynamo framework and loads new checkpoints into a live deployment. In testing, CoreWeave says, it improved model reload latency by 15x compared with a baseline configuration.
Chowdhary explained the practical effect: training keeps moving quickly while inference scales independently in parallel, instead of one waiting for the other.
Forge: one platform, free entry point
RL Rollouts and the inference optimizations now live inside Forge. The platform links serving, observability, post-training and evaluation, and it is free to start. Paid tiers add more capabilities.
That pricing model is aimed squarely at individual developers, not just large enterprise accounts. Chowdhary framed it as an accessibility move: someone signing up alone should get the same performance and reliability as a bigger customer, without compromises.
The Bigger Picture
For readers tracking AI infrastructure, Forge is a signal of where specialized GPU clouds think the money is going. If inference grows from a side workload into the main one, as the healthcare example suggests, then providers that only sell compute risk becoming interchangeable. Bundling software, tooling and evaluation is one way to avoid that.
The move also fits a broader pattern. CoreWeave has recently appeared in partnerships such as its work with IBM to co-design controls for agent workloads, and rivals like Lambda are raising capital ahead of a planned IPO. Meanwhile, startups such as Seismora are building control planes that route AI workloads across devices and clouds. The layer between raw hardware and the finished application is getting crowded.
The RL Rollouts feature is perhaps the more telling detail. It suggests that agentic model training is mature enough for infrastructure vendors to build features around its specific pain points. The 15x figure, though, comes from CoreWeave's own testing against a baseline it chose, and the capability is still in preview. Independent results would help.
It is worth watching three things next: whether Forge's free tier actually draws individual developers away from larger clouds, how the paid tiers are priced, and whether CoreWeave's commitment to open source tools holds as the platform becomes a bigger part of its business.
