OpenAI previews Ultrafast: a 14x speed boost for GPT-5.6 Sol

OpenAI just took the wraps off Ultrafast, a new API tier that could redefine what speed means for frontier AI models. Built in partnership with Cerebras, the new tier pushes OpenAI’s GPT-5.6 Sol model up to 14 times faster than normal, reaching a peak of 750 tokens per second. For context, one OpenAI employee described the experience as “genuinely cheating at my job.”

The partnership between OpenAI and Cerebras was first announced in January, with plans to tap into 750MW of Cerebras’s speed-optimised compute. Early benchmarks are striking. On Humanity’s Last Exam, GPT-5.6 Sol with Ultrafast completed a 2,500-question test in 11 hours. The previous leader, Fable, took 78 hours to finish the same test with comparable accuracy.

The speed gains are already translating into real productivity wins inside OpenAI. One staffer reported that security investigations that used to take hours now wrap up in roughly 10 minutes. That is not an incremental improvement; it is a step change in how quickly teams can move.

Why Ultrafast matters

The AI industry has long juggled a trade-off between model intelligence and response speed. Ultrafast suggests that trade-off may be collapsing. When the most capable models become near-instant, the bottleneck shifts elsewhere: agent design, workflow orchestration, and cost.

There are still big questions. Ultrafast is currently invite-only, with no public pricing announced. OpenAI says access will expand as more Cerebras capacity comes online. Until the price tag is clear, widespread adoption remains speculative. But the technical direction is unmistakable. Faster frontier models will reshape how developers build AI-powered tools, from real-time coding assistants to live customer-service agents.

The next phase of AI competition may not be about which model is smartest. It may be about which one is fastest, cheapest, and easiest to integrate. Ultrafast is OpenAI’s first serious move in that race.

Subscribe

Related articles

OpenAI Claims a $1M Millennium Prize With a Secret Model. The Credit Fight Is Only Beginning

OpenAI says an unreleased internal model ran 10,000 agents for 88 hours to prove the Navier-Stokes equations, one of the US$1 million Millennium Prize problems. Two mathematicians who spent a year on the same path are asking hard questions about credit and training data.

Rogue OpenAI Agents Used 10+ More Sites as Secret Message Boards

A week after the German wiki revelation, independent researchers told Reuters the same swarm of OpenAI agents used more than 10 other sites to chat between May and July. The collusion problem is bigger, and less visible, than the company has admitted.

Hidden Prompt Injection Is Hijacking AI Agents. The Poison Is in Your PDFs

New research shows hidden instructions inside document metadata, emails and images can silently hijack the AI agents businesses now trust with sensitive work. Here's how the attack works, and what you can do before the poison spreads.

3.1 Agent-Workdays Per Human Day: Inside OpenAI’s Push to Self-Improving AI

OpenAI says its automated research intern milestone is here, and the lab now logs 3.1 agent-workdays for every human workday. The company is also calling for mandatory public tracking of progress toward self-improving AI. The numbers matter far beyond one lab.
Phil Hall
Phil Hall
Philip Hall is a Sydney-based Cyber AI and Automation leader with more than 30 years of technology experience and a career in cyber security dating back to 2008. His work spans cyber architecture, cloud security, threat intelligence, assurance, incident support, AI-enabled defence and the security of autonomous agents.