OpenAI Launches ‘Ultrafast’ Mode for GPT-5.6 Sol, Boosting Inference Speed by 14x

Close up view of the ChatGPT mobile application icon on a smartphone screen

Quick Read

  • OpenAI introduced 'Ultrafast' mode for GPT-5.6 Sol, reaching 14x faster speeds.
  • The service is powered by Cerebras's wafer-scale engine technology.
  • Ultrafast generates up to 750 output tokens per second.
  • The mode is currently in a limited preview for select API customers.
  • It is designed for time-sensitive tasks like cyberattack response and financial modeling.

A New Tier for Real-Time Frontier Intelligence

OpenAI has officially unveiled “Ultrafast” mode, a specialized service tier designed to significantly accelerate the inference speed of its flagship model, GPT-5.6 Sol. Launched in partnership with chipmaker Cerebras, the new mode is capable of delivering up to 750 output tokens per second, effectively operating at 14 times the speed of the standard processing tier. The initiative aims to bridge the long-standing gap between high-level reasoning capabilities and the latency requirements of real-time enterprise applications.

Historically, AI developers have faced a binary choice: utilize smaller, specialized models for speed, or deploy larger, frontier-level models at the cost of slower response times. By utilizing Cerebras’s wafer-scale engine architecture, OpenAI claims that GPT-5.6 Sol can now handle complex, mission-critical tasks without sacrificing the reasoning quality associated with its frontier models.

Technical Architecture and Benchmarking

The speed breakthrough is attributed to the unique hardware approach taken by Cerebras. Traditional GPU-based systems often struggle with memory bandwidth bottlenecks, where model weights must be repeatedly moved between on-chip memory and off-chip storage. Cerebras mitigates this by packing 44 GB of SRAM directly onto wafer-sized chips, keeping weights on-chip and allowing tokens to flow through pipelined model layers without interruption.

In comparative testing, Cerebras evaluated GPT-5.6 Sol on Ultrafast mode against other industry benchmarks using “Humanity’s Last Exam” (HLE)—a dataset consisting of 2,500 complex questions spanning fields such as economics, chemistry, and literature. The results showed that GPT-5.6 Sol completed the exam in 11 hours and 11 minutes, while the comparison model, Claude Fable 5, required over 78 hours to reach similar conclusions.

Enterprise Implications

The deployment of Ultrafast is expected to impact high-stakes corporate workflows where response time is a direct performance metric. OpenAI and Cerebras highlighted several key sectors for early adoption:

  • Incident Response: Security teams can leverage the speed to identify and contain cyberattacks in real-time, potentially preventing catastrophic losses.
  • Financial Analysis: Rapid processing enables faster responses to fluctuating market conditions and complex data modeling.
  • Customer Support: Organizations can deploy sophisticated, high-reasoning agents that provide instant, high-quality resolution to customer queries.
  • Engineering and Legal: The speedup allows for the rapid drafting and analysis of complex documents, such as legal briefs and engineering reports, without the wait times typically associated with large LLMs.

Availability and Future Outlook

As of August 13, 2026, the Ultrafast mode is available in a limited preview within the OpenAI API. Access is currently restricted to a select group of customers, with plans to expand capacity as the service matures. The collaboration between OpenAI and Cerebras signals a broader industry trend toward optimizing hardware-software integration to push the boundaries of what is possible with large-scale agentic workflows.

|
Creator:Azat TV Editorial

LATEST NEWS