CoreWeave to Offer NVIDIA Vera, First CPU Built for AI Agents
CoreWeave will expand its compute portfolio by offering the NVIDIA Vera CPU, the first CPU designed for AI agents. This expansion aims to meet the needs of new workloads, particularly agentic AI, which requires significant CPU resources for tasks surrounding AI model training and reasoning. CoreWeave emphasizes that Vera will run on their bare-metal platform, offering enhanced performance and scalability for AI infrastructure.
Story updates
01:03:49 PM UTC
SquawkNews
For best results when printing this announcement, please click on link below: CoreWeave Delivers NVIDIA Vera Rubin NVL72 Performance at Production Scale, Starting With Cognition CoreWeave maintains industry-leading AI cloud performance across multiple compute generations with its software platform CoreWeave Inc. (Nasdaq: CRWV), The Essential Cloud for AI™, today announced the availability of NVIDIA Vera Rubin NVL72 on CoreWeave, with Cognition as the first customer anywhere running production workloads on the system. Customers, like Cognition, run the system under the same operating model and tooling as their existing NVIDIA GB200 NVL72 and GB300 NVL72 fleets, with performance engineering from CoreWeave’s team. The news was shared during Fully Connected ( ), CoreWeave’s AI cloud conference, which brings together more than 4,500 customers, partners, developers and AI leaders to share how they are building and running AI in production. This press release features multimedia. View the full release here: NVIDIA Vera Rubin NVL72 systems deployed on CoreWeave Cloud for production AI workloads. “Bringing up NVIDIA Vera Rubin NVL72 so quickly, and having a customer already seeing performance gains within days, is the payoff from years of engineering our platform across GPU generations,” said Chen Goldberg, executive vice president of product & engineering at CoreWeave. “With customers like Cognition, that investment shows up in the ability to get production workloads running within days. When it comes to agentic tasks, long contexts, repeated model calls and thousands of concurrent tasks put pressure on the entire platform. Our job is to make compute, networking and software work as a single system, so customers can build increasingly complex agents without taking on the infrastructure complexity themselves.” Cognition is the first customer in production with Vera Rubin NVL72 Cognition, the applied AI lab behind Devin, runs training, reinforcement learning and production inference for Devin on CoreWeave. The company scaled from bridge capacity to thousands of GPUs for training and inference in less than nine months. Cognition worked with CoreWeave to stand up a Vera Rubin NVL72 cluster in early September, and Cognition’s own engineers ran the first customer-executed Vera Rubin inference benchmark, measured against a GB200 NVL72 cluster baseline. “Agentic coding is an unforgiving workload that requires long contexts, high concurrency and rapid reasoning,” said Silas Alberti, SVP research & founding team at Cognition. “By deploying the NVIDIA Vera Rubin NVL72 on CoreWeave, our engineers are seeing up to a 4.8 times increase in total token throughput for SWE-2 inference workloads. For an agentic workload where every step waits on the last one, that compounds into real work Devin gets done. CoreWeave continues to deliver the bleeding-edge rack-scale acceleration we need to push the boundaries of AI.” Cognition’s engineers benchmarked Vera Rubin I