OpenAI's GPT-6 Astra Ultrafast Runs on NVIDIA Blackwell GPUs

OpenAI's GPT-6 Astra Ultrafast is available now in the API and to eligible ChatGPT Work and Codex users. It runs on NVIDIA Blackwell GPUs, and NVIDIA says it generates tokens up to 8x faster than in…

Oct 2, 2026
•
3 min read
Technobezz
OpenAI's GPT-6 Astra Ultrafast Runs on NVIDIA Blackwell GPUs

Don't Miss the Good Stuff

Get tech news that matters delivered to your inbox.

OpenAI's GPT-6 Astra Ultrafast is available now through the OpenAI API, and the model also reaches eligible ChatGPT Work and Codex users. It runs on NVIDIA Blackwell GPUs, and NVIDIA says the setup delivers up to 8x faster token generation than Astra Standard mode.

The speedup comes from inference optimizations that lean on the Blackwell architecture, according to the announcement. NVIDIA frames the benefit in practical terms: faster generation can shorten the edit-test-debug cycles coding agents run through, can trim the time a model spends answering between tool calls, and can leave interactive applications feeling more responsive.

Read more: Anthropic Brings Claude AI Models to Microsoft Foundry on NVIDIA Blackwell Hardware

Developers can reach Ultrafast through the API today, and a guide covers access, pricing and implementation details. The announcement does not disclose price figures, token limits, context window sizes, benchmark results, or a regional availability breakdown.

OpenAI uses its own models to refine inference software on NVIDIA GPUs, and NVIDIA says that ongoing optimization can speed responses while raising infrastructure productivity. The programmable NVIDIA platform also lets teams reuse the same infrastructure across training, inference and reinforcement learning, which NVIDIA says helps repurpose compute as demand shifts and avoids overprovisioning.

The launch extends a GPT-6 family that already reached GitHub Copilot. Copilot users can pick between GPT-6 Sol, aimed at interactive and agentic coding, and GPT-6 Luna, the cheapest member of the family, both billed on usage. That rollout was gradual, so not every eligible user saw the models at once, and it spanned editors and surfaces including Visual Studio Code, JetBrains, Xcode, Copilot CLI, github.com and GitHub Mobile.

Copilot Business and Enterprise administrators control those models through a model policy, with default enablement switching new models on automatically and admins able to change the global default or disable each model. That earlier wave also brought Claude Opus 5.5 to Copilot the same day and Grok 4.7 a day earlier at provider list pricing. GitHub plans to deprecate several models on October 19, among them Gemini 3.7 Flash, GPT-5.5, GPT-5.4, GPT-5.4 mini, GPT-5 mini and Grok 4.5, across Chat, inline edits, ask, agent and completions.

Philippe Tillet, OpenAI's inference lead, pointed to NVIDIA's tooling and documentation, saying "NVIDIA’s deep investment in tooling and documentation has enabled us to make our models exceptionally good at programming Blackwell and Rubin GPUs".