On August 13 OpenAI unveiled Ultrafast — a mode in which its flagship GPT-5.6 Sol answers up to 14 times faster than usual. It is not a new model and not a stripped-down version: the intelligence is the same, only the token output speed changes. The mode is still a limited preview in the API, but Sol itself already runs in GPTunneL — we covered what it can do in a separate article.
What Ultrafast actually is
Ultrafast is not a model but a service tier sitting next to the regular Standard one. Your request goes to the very same GPT-5.6 Sol, the answer is identical in quality, but the tokens arrive several times faster.
The old trade-off was uncomfortable: if you needed speed, you picked a smaller, simpler model — and accepted weaker answers. That is exactly the fork OpenAI is removing. The company puts it bluntly: until now, real-time speed meant switching to a smaller or more specialized model.
Don't confuse Ultrafast with Ultra, the mode Sol has had since its announcement. Ultra is the opposite — it's about thinking longer: the model spins up several internal sub-agents and works on different parts of a hard task in parallel. Ultrafast is about an instant answer at the same level of intelligence. Two ends of one scale.
The key numbers
- up to 750 output tokens per second — the mode's peak speed;
- up to 14× versus Standard: Sol's baseline is around 53 tokens per second;
- no quality trade-off — the model is not trimmed or swapped for a distilled version.
To feel the difference: 53 tokens per second is roughly the pace of someone reading aloud, slow enough to follow with your eyes. At 750 tokens per second a paragraph is fully on screen before you finish its first line.
Where the speed comes from
Ultrafast does not run on conventional GPUs but on Cerebras hardware — the Wafer-Scale Engine architecture. The chip is the size of an entire silicon wafer and carries 44 GB of SRAM right on board.
The point is where the model weights live. On ordinary accelerators they have to be shuttled between on-chip and external memory, and it is that bandwidth — not the compute — that caps inference speed for frontier models. Cerebras keeps the weights on the chip itself, the bottleneck disappears, and that is where the multiple-times gap comes from.

What the benchmarks show
The main claim is that the speed was not bought with quality.
- Humanity's Last Exam. 2,500 PhD-level questions across chemistry, economics, literature and other fields. Sol on Ultrafast worked through the whole set in 11 hours 11 minutes; Claude Fable 5 needed 78 hours 27 minutes on the same set — more than three days of continuous compute. Roughly a sevenfold difference at comparable accuracy.
- GDP-Val, a suite modelling real economically valuable knowledge work: an end-to-end speedup of 5.6× with no quality degradation.
- Per Cerebras' own measurements, Sol on Ultrafast is 11 times faster than Fable 5 and 5 times faster than Opus 4.8 running in its own fast mode.
One honest caveat: every number above comes from OpenAI and Cerebras. There are no independent measurements yet — those will arrive once the mode opens up and third-party labs get their hands on it.
Who actually needs that speed
OpenAI names the scenarios where seconds decide the outcome:
- incident response — the model reads logs and traces while the outage is still happening, not after;
- voice assistants and live support — in a spoken conversation a few seconds of silence kills the dialogue;
- financial research — market analysis feels like a live conversation rather than send-and-wait;
- commerce — recommendations and matching while the shopper is still on the page;
- developer agents — long chains of actions stop being bottlenecked by generation speed;
- research. The most visible effect: runs that used to be queued overnight now fit inside a working day, so you get several iterations instead of one per day.
Among the companies already testing the mode OpenAI names Jane Street, Podium, Basis and Rogo — finance, support and internal product teams.
Sol, Terra or Luna — which one to pick
The GPT-5.6 family has three models, and Ultrafast does not settle the choice between them: it is about the speed of one specific model, not about price.
| Model | What it's for | Price |
|---|---|---|
| GPT-5.6 Sol | The flagship: hard code, research, long agentic chains | The priciest of the three |
| GPT-5.6 Terra | The workhorse: nearly the same quality, noticeably cheaper | Roughly two to three times cheaper than Sol |
| GPT-5.6 Luna | High-volume requests, automation, heavy load | The cheapest — dozens of times cheaper than Sol |
The practical rule is simple: push drafts, classification and repetitive jobs through Luna, everyday work through Terra, and save Sol for the tasks that genuinely need the ceiling. Current figures for all three are on the pricing page — they are recalculated from the platform catalogue. A detailed breakdown of the differences is in our GPT-5.6 Sol article.
When will Ultrafast open to everyone
For now it is a limited preview: a small group of companies has access, and OpenAI promises to widen it as capacity grows. You can join the queue by describing your workload, latency requirements and expected volumes.
The price of the mode has not been announced — neither in the launch post nor in the docs. It is reasonable to expect speed to cost more than the standard tier, but for now that is an expectation, not a fact.
Where to try GPT-5.6 Sol today
Ultrafast is behind a closed preview, but the flagship is available right now: open GPT-5.6 Sol in GPTunneL — no subscription, pay per use. Terra and Luna sit in the same chat, alongside Claude, Gemini, DeepSeek and the rest of the catalogue, so you can switch between them in one window on a single balance.
FAQ
Is Ultrafast a new model? No. It is a service tier for the existing GPT-5.6 Sol. Same weights, same answer quality, same capabilities — only the output speed differs.
Can I turn Ultrafast on in GPTunneL? Not yet: the mode is in a closed OpenAI API preview and has not been opened publicly. We'll write about it in the blog as soon as it is. Sol in standard mode already works here.
How much faster is Ultrafast really? Per OpenAI, up to 750 tokens per second versus around 53 in standard mode — up to 14 times. That "up to" matters: peak speed and the average across real tasks are not the same thing.
Does answer quality drop in the fast mode? OpenAI and Cerebras say it does not, and back that up with HLE and GDP-Val results. Independent measurements are not available yet.
How does Sol differ from Terra and Luna? Sol is the flagship for hard tasks, Terra balances quality and cost, Luna is the fastest and cheapest for high volume. Details are in our GPT-5.6 Sol breakdown.
Try it yourself
Ultrafast is a signal of where the industry is heading: speed is no longer the price you pay for intelligence. While the mode is being handed out to a select few, OpenAI's strongest model is already available with no waiting — launch GPT-5.6 Sol in GPTunneL and compare it against Terra and Luna on your own tasks.
New models and modes land here the day their API opens — follow the news in our blog.



