Back to feed
News Story
APriority81
机器之心
3 sources

OpenAI and Cerebras Preview GPT-5.6 Sol Ultrafast Mode with 14x Speed Boost

OpenAI, in partnership with AI chip maker Cerebras, previewed the Ultrafast Mode for its flagship model GPT-5.6 Sol, achieving output speeds up to 750 tokens per second—14 times faster than the standard mode without quality loss. The mode leverages Cerebras' wafer-scale hardware and is rolling out via the OpenAI API to select customers in a limited preview.

SynthePulse Insight · AI deep readingMembers

GPT-5.6 Sol Ultrafast Mode: The Architecture Revolution and Business Logic Behind 14x Acceleration

Version 1 · 3 sources

OpenAI, in partnership with Cerebras, has introduced the Ultrafast mode for GPT-5.6 Sol, achieving output speeds of up to 750 tokens per second—14 times faster than the standard mode. This breakthrough not only relies on the on-chip memory of wafer-scale chips but also signals a new trend where AI inference speed becomes an independent commercial dimension.

  • OpenAI previews Ultrafast mode for GPT-5.6 Sol, delivering up to 750 tokens/s output—14x faster than standard mode with no quality loss.
  • The mode is powered by Cerebras' wafer-scale chip WSE-3, whose 44 GB on-chip SRAM keeps model parameters resident, avoiding memory bandwidth bottlenecks.
  • On the HLE benchmark, Ultrafast mode completed all questions in 11 hours 11 minutes, while Claude Fable 5 took 78 hours 27 minutes—roughly 7 times faster.
Open section navigation14x Acceleration: The Performance Leap Behind the Numbers

14x Acceleration: The Performance Leap Behind the Numbers

On August 14, early in the morning, OpenAI, in collaboration with AI chip maker Cerebras, officially previewed a new service tier for its flagship model GPT-5.6 Sol: "Ultrafast Mode." In this mode, output speeds can reach up to 750 tokens per second, a maximum 14x improvement over the Standard mode's inference baseline of about 53 tokens per second, with no reduction in quality. In horizontal comparison, accelerated GPT-5.6 Sol is 11 times faster than Fable 5 and 5 times faster than Opus 4.8 in Fast mode.

In Cerebras' blog, engineers described tests on the "Humanity's Last Exam" (HLE). HLE contains 2,500 questions, typically solvable only by PhDs. GPT-5.6 Sol in Ultrafast mode answered all questions in just 11 hours 11 minutes, while Claude Fable 5 required 78 hours 27 minutes—about 7 times faster, with similar accuracy. On the GDP-Val benchmark, Ultrafast achieved a 5.6x end-to-end speedup with no impact on quality.

These numbers indicate that Ultrafast mode not only boosts throughput but also significantly shortens completion time for complex tasks, enabling real-time interactive workflows.

Free for now

Read the full analysis

4 more sections of analysis, plus the full takeaway

Loading

Credibility boundary

This article's information primarily comes from reports by Jiqizhixin, THE DECODER, and TechCrunch, as well as official blogs from OpenAI and Cerebras. All performance data (such as 750 tokens/s, 14x speedup, HLE times) are from official releases or media reports but have not been independently verified by third parties. Cerebras' hardware specifications (4 trillion transistors, 125 petaflops, 44 GB SRAM) come from its official blog. Details on pricing and partnerships come from THE DECODER's report, which is second-hand information.

Primary report

机器之心

Primary source

Same-event coverage

Also covered by 2 sources