Cognition is achieving 4.8 times higher token throughput for its Devin AI software engineer by running training, reinforcement learning and production inference on newly available NVIDIA Vera Rubin NVL72 systems at CoreWeave, the company says. Cognition, the applied AI lab behind the Devin AI software engineer, is the first customer running production workloads on Vera Rubin.
“Having all of it on one platform, with NVIDIA and CoreWeave engineers who work the hard problems alongside ours, matters more to us than any single spec,” says Silas Alberti, founding team at Cognition. CoreWeave has launched CoreWeave Forge, a connected environment for developing and refining AI models on NVIDIA accelerated computing.
NVIDIA Vera Rubin NVL72 Accelerates Cognition’s AI Workloads
Cognition is realizing a 4.8x performance gain. This performance gain stems from a production cluster established for Cognition in a matter of days, a rapid deployment facilitated by the close collaboration between NVIDIA and CoreWeave across the entire technology stack. Capacity is managed through CoreWeave’s Kubernetes Service, SUNK, Mission Control, Sandboxes and Inference offerings, allowing for scalable and efficient operation of demanding AI workloads.
The speed of deployment highlights a shift towards integrated infrastructure solutions designed to accelerate the AI lifecycle from model development to real-world application. CoreWeave’s commitment to long-term hardware use is also apparent, with NVIDIA V100 GPUs, based on the Volta architecture, still actively serving customer workloads nearly a decade after their initial release.
This sustained productivity demonstrates the durability and value proposition of NVIDIA’s accelerated computing platform, allowing CoreWeave to continue using existing investments while simultaneously introducing new technology, according to the company. “NVIDIA accelerated computing delivers value across generations,” said Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA, emphasizing the longevity and adaptability of their infrastructure solutions. The ability to transition smoothly between GPU generations minimizes disruption and maximizes return on investment for customers like Cognition.
The complexity of agentic coding, characterized by long context windows, high concurrency and substantial token volumes, demands a cost-effective approach to processing. “Agentic coding is a complex workload: long contexts, high concurrency and token volumes where cost per token decides what we can ship,” explained a source within Cognition, highlighting the importance of optimizing performance and efficiency. CoreWeave’s integration of the NVIDIA Vera Rubin NVL72, alongside the forthcoming NVIDIA Vera CPU, aims to address these challenges by providing a purpose-built cloud environment tailored for the unique requirements of AI agents.
NVIDIA accelerated computing delivers value across generations.
Ian Buck, vice president of hyperscale and high-performance computing at NVIDIA
CoreWeave Forge Unifies AI Training, Evaluation, and Production
CoreWeave is streamlining the AI development process with the launch of CoreWeave Forge, a platform designed to unify training, evaluation and production deployment of AI models. The system integrates tools like Weights & Biases, OpenPipe and the marimo notebook project into a single environment, aiming to reduce data loss during transitions between development stages. This consolidation allows for a continuous feedback loop where production data directly informs subsequent training runs, accelerating model refinement.
The platform’s impact is already demonstrable, with Cognition achieving 4.8x performance boost benchmarked against a GB200 NVL72 baseline using a real-world software engineering workload, highlighting the efficiency gains possible with the integrated infrastructure. CoreWeave’s approach also improves failure detection by 20 percent and reduces issue resolution costs by half, converting production agent data into actionable insights.
CoreWeave Sandboxes, now generally available, further enhance the development process by providing isolated environments for running agents, tool calls, reinforcement learning and evaluations, the firm reports. These sandboxes operate on either serverless infrastructure or existing training infrastructure, ensuring a consistent and secure environment for every agent interaction. The company also reported a 1.7x performance gain on Vera CPU across all passing tasks on Terminal-Bench, demonstrating the benefits of the new hardware in broader testing scenarios. CoreWeave Forge aims to close the loop between production and training, enabling faster iteration and improved AI model performance.
NVIDIA Vera CPU Enables 3x Faster Agent Sandbox Startups
CoreWeave announced availability of NVIDIA Vera Rubin NVL72 on CoreWeave Cloud, making it one of the first cloud providers to deliver the platform to customers. Early-access customers can put the performance of NVIDIA’s full-stack AI factory platform to work quickly on CoreWeave Cloud. In days, CoreWeave stood up a production Vera Rubin cluster for Cognition, achieved as a result of the codesign and collaboration between NVIDIA and CoreWeave up and down the stack, from infrastructure to tokens served, the company states.
CoreWeave will also offer NVIDIA Vera, the first CPU built for AI agents. The infrastructure is designed to address the demands of agentic AI, which places significant pressure on computing resources both for serving agents and for improving them through post-training processes.
CoreWeave uses Spectrum-X Ethernet switches and BlueField-4 DPUs to ensure secure, high-performance and low-latency communication between these isolated agent environments. This integrated approach allows for RL Rollouts, a technique that loads new checkpoints into live deployments, enabling continuous reinforcement learning without requiring redeployments and maintaining inference efficiency. Cognition was the first customer to receive a production Vera Rubin cluster from CoreWeave, demonstrating the speed with which the platform can be made available to early-access users, by the company’s account.
Companies like Canva, Capital One, and MasterClass are among the first to build on this new infrastructure, using NVIDIA Nemotron open models to customize and deploy reasoning and multimodal models for agentic workflows. The co-engineered NVIDIA and CoreWeave platform is designed to accelerate the transition from prototype to production for AI labs, AI-native companies and global enterprises.




See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
