OpenAI just took the wraps off Ultrafast, a new API tier that could redefine what speed means for frontier AI models. Built in partnership with Cerebras, the new tier pushes OpenAI’s GPT-5.6 Sol model up to 14 times faster than normal, reaching a peak of 750 tokens per second. For context, one OpenAI employee described the experience as “genuinely cheating at my job.”
The partnership between OpenAI and Cerebras was first announced in January, with plans to tap into 750MW of Cerebras’s speed-optimised compute. Early benchmarks are striking. On Humanity’s Last Exam, GPT-5.6 Sol with Ultrafast completed a 2,500-question test in 11 hours. The previous leader, Fable, took 78 hours to finish the same test with comparable accuracy.
The speed gains are already translating into real productivity wins inside OpenAI. One staffer reported that security investigations that used to take hours now wrap up in roughly 10 minutes. That is not an incremental improvement; it is a step change in how quickly teams can move.
Why Ultrafast matters
The AI industry has long juggled a trade-off between model intelligence and response speed. Ultrafast suggests that trade-off may be collapsing. When the most capable models become near-instant, the bottleneck shifts elsewhere: agent design, workflow orchestration, and cost.
There are still big questions. Ultrafast is currently invite-only, with no public pricing announced. OpenAI says access will expand as more Cerebras capacity comes online. Until the price tag is clear, widespread adoption remains speculative. But the technical direction is unmistakable. Faster frontier models will reshape how developers build AI-powered tools, from real-time coding assistants to live customer-service agents.
The next phase of AI competition may not be about which model is smartest. It may be about which one is fastest, cheapest, and easiest to integrate. Ultrafast is OpenAI’s first serious move in that race.