OpenAI has officially launched the highly anticipated GPT-5.6Sol "Ultra-Fast" mode. Prior to this, the version was internally previewed in July for certain specific customers. Now, the output speed of this mode can reach up to 750 tokens per second. Without switching to a smaller or lower-performance model, the overall processing speed is up to 14 times faster than the standard mode.

This technological breakthrough is mainly supported by Cerebras' wafer-scale engine architecture. This architecture includes 44GB of on-chip static random access memory, successfully breaking the common data transfer bottlenecks found in traditional GPU systems. According to relevant test data, its operating speed not only significantly surpasses competitors but also processed all 2,500 questions in the "Ultimate Human Exam" in about 11 hours, demonstrating extremely high computing efficiency.
Currently, this new mode is very suitable for high-intensity workflows such as incident response, financial research, real-time customer support, programming, and complex research. Given that computing resources are still limited, OpenAI will gradually assess and grant access based on the fit of the customer's workload and the availability of computing power.
