On August 14, early in the morning, OpenAI, in collaboration with AI chip maker Cerebras, officially previewed a new service tier for its flagship model GPT-5.6 Sol: "Ultrafast Mode." In this mode, output speeds can reach up to 750 tokens per second, a maximum 14x improvement over the Standard mode's inference baseline of about 53 tokens per second, with no reduction in quality. In horizontal comparison, accelerated GPT-5.6 Sol is 11 times faster than Fable 5 and 5 times faster than Opus 4.8 in Fast mode.
In Cerebras' blog, engineers described tests on the "Humanity's Last Exam" (HLE). HLE contains 2,500 questions, typically solvable only by PhDs. GPT-5.6 Sol in Ultrafast mode answered all questions in just 11 hours 11 minutes, while Claude Fable 5 required 78 hours 27 minutes—about 7 times faster, with similar accuracy. On the GDP-Val benchmark, Ultrafast achieved a 5.6x end-to-end speedup with no impact on quality.
These numbers indicate that Ultrafast mode not only boosts throughput but also significantly shortens completion time for complex tasks, enabling real-time interactive workflows.