Back to feed
News Story
APriority74
OpenAI Blog
1 sources

OpenAI Previews Ultrafast Mode: GPT-5.6 Sol at Up to 14x Speed

OpenAI has previewed a new API service tier called Ultrafast, powered by Cerebras, that runs GPT-5.6 Sol up to 14 times faster, delivering up to 750 output tokens per second. This new tier aims to provide developers with faster AI inference, potentially impacting applications that rely on real-time responses.

SynthePulse Insight · AI deep reading

OpenAI Previews Ultrafast: GPT-5.6 Sol 14x Faster, Powered by Cerebras Hardware

Version 1 · 1 source

On August 13, 2026, OpenAI issued a preview announcement introducing a new API service tier called Ultrafast, claiming to boost GPT-5.6 Sol inference speed by 14x, with output rates up to 750 tokens per second. The service is powered by Cerebras.

  • OpenAI previews Ultrafast, a new API service tier designed for GPT-5.6 Sol.
  • Ultrafast claims up to 14x speed improvement and 750 output tokens per second.
  • The service is powered by Cerebras hardware, but specific technical details are undisclosed.
  • Currently in preview; official launch date and pricing not announced.
Open section navigationAnnouncement Core: Speed and Hardware

Announcement Core: Speed and Hardware

On August 13, 2026, OpenAI published a preview announcement via its official blog, introducing a new API service tier called Ultrafast. The service is designed specifically for the GPT-5.6 Sol model, claiming to boost its runtime speed by 14x, with output rates reaching 750 tokens per second.

The announcement explicitly states that Ultrafast is powered by Cerebras. Cerebras is a company known for manufacturing large wafer-scale chips, but the announcement does not disclose specific hardware configurations or optimization techniques.

Performance Metrics Interpretation

The 14x speed improvement and 750 output tokens per second are the core numbers in the announcement. These metrics may be based on specific test conditions, but the announcement does not provide benchmark details or comparison baselines.

Notably, these figures come from OpenAI's official announcement and are self-reported, without independent verification.

Service Tier and Availability

Ultrafast is described as an 'API service tier,' suggesting it may be offered as a paid option separate from the standard API, but the announcement does not mention pricing, official release date, or regional availability.

Currently, the service is in preview; developers may gain access through application or invitation, but the announcement does not clarify the process.

Industry Context and Potential Impact

This release comes amid intensifying competition in AI inference speed. OpenAI's collaboration with Cerebras may aim to meet the demand for low-latency, high-throughput inference, especially in real-time application scenarios.

However, the announcement does not mention any customer cases or actual deployments, so its real-world performance remains to be seen.

Credibility boundary

This report is based on a single source: OpenAI's official blog. All performance figures and collaboration information come from that announcement and are self-reported, without independent verification.

Insight takeaway

OpenAI previews Ultrafast, claiming a 14x speed improvement for GPT-5.6 Sol inference, with output rates of 750 tokens per second, powered by Cerebras. However, the announcement lacks technical details and independent verification, so actual performance and availability remain to be seen.

Primary report

OpenAI Blog

Primary source