OpenAI’s GPT-6 Astra Ultrafast now generates tokens up to eight times faster than the Astra Standard mode, a performance leap enabled by NVIDIA Blackwell GPUs, the company says. The accelerated model impacts applications from code generation to interactive tools, shortening critical edit-test-debug cycles for developers. “NVIDIA’s deep investment in tooling and documentation has enabled us to make our models exceptionally good at programming Blackwell and Rubin GPUs,” says Philippe Tillet, inference lead at OpenAI. GPT-6 Astra Ultrafast is currently accessible through the OpenAI API and to eligible ChatGPT Work and Codex users.
Blackwell GPUs Enable 8x Faster GPT-6 Astra Ultrafast Token Generation
This acceleration stems from inference optimizations developed by OpenAI, specifically using the programmability of NVIDIA GPUs. OpenAI’s models were employed to refine the inference software running on Blackwell, continually improving performance and maximizing the platform’s capabilities, according to the company. The benefits of faster processing are most pronounced in iterative workflows where agents repeatedly write code, use tools and evaluate results. OpenAI’s Astra Ultrafast brings these capabilities into time-sensitive loops, providing more useful model outputs when developers require them.

OpenAI’s strategy extends beyond initial deployment. The company is actively using its own models to optimize inference on NVIDIA GPUs. This programmable NVIDIA platform also offers flexibility, allowing developers and researchers to reuse infrastructure across training, inference and reinforcement learning as models evolve, improving resource use and avoiding overprovisioning.
NVIDIA’s deep investment in tooling and documentation has enabled us to make our models exceptionally good at programming Blackwell and Rubin GPUs.
Philippe Tillet, inference lead at OpenAI




See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
