infra
Previewing Ultrafast mode: GPT-5.6 Sol at up to 14X the speed
August 13, 2026
OpenAI is previewing Ultrafast, a new API service tier that runs GPT-5.6 Sol up to 14× faster and, with Cerebras, can deliver up to 750 output tokens per second. The speedup matters for latency-sensitive applications and shows how specialized hardware is being used to push frontier model inference rates much higher.
Preview Ultrafast, a new OpenAI API service tier that runs GPT-5.6 Sol up to 14× faster. Powered by Cerebras, it delivers up to 750 output tokens per second.
Source: openai.com