Peak output throughput — frontier intelligence at reading‑defying speed.
New speed class · Now in limited preview
Ultrafast mode. GPT‑5.6 Sol at up to 14× the speed.
Frontier intelligence at 750 output tokens per second — the same GPT‑5.6 Sol, up to fourteen times faster than Standard processing, launching first in the OpenAI API. Powered by Cerebras.
01 — THE TRADE‑OFF IS OVER
Real‑time speed used to mean a smaller, dumber model.
Until now, getting real‑time speed typically meant choosing a smaller or more specialized model — trading away the intelligence your hardest problems need. Ultrafast points in a new direction: more useful work per second.
When speed no longer requires giving up intelligence, AI moves into the most time‑sensitive parts of a business — and new kinds of work become possible. The same GPT‑5.6 Sol frontier model, running up to 14× faster, on ultra‑low‑latency hardware.
Specifications
Versus Standard processing on GPT‑5.6 Sol — roughly 54 tokens per second.
Our most intelligent model — unchanged. No smaller stand‑in, no distilled substitute.
Launching first in the OpenAI API, in limited preview for select customers.
02 — THE CENTERPIECE
Same prompt. Same brain. Different clock.
Watch GPT‑5.6 Sol answer the exact same request in Ultrafast and Standard modes, token by token, at published rates. This is what fourteen times looks like.
Ultrafast lane idle
Standard lane idle
Simulated locally at published rates — up to 750 tok/s vs ~54 tok/s. Answers are illustrative, scripted output.
03 — THE THROTTLE
Push it anywhere from 54 to 750 tok/s.
Pick a task or write your own, drag the throttle, and watch the same intelligence answer at any speed in the new class — with Standard processing along for comparison.
Pick a task and hit Generate — the answer streams at the throttle speed.
Playground idle
Demo simulation — custom prompts get a scripted illustrative answer, streamed character‑for‑token at your chosen rate.
04 — WHERE SPEED WINS
The most time‑sensitive work in your business, unblocked.
Five workflows where an order‑of‑magnitude jump in speed changes what the product even is.
Incident response & reliability
When a critical system fails, Sol reads the application logs, the deploy diff, and the on‑call thread while the outage is still unfolding — and hands back a ranked hypothesis with a candidate fix before the first bridge call warms up. Mean time to understand collapses; engineers stay in charge of judgment and deployment.
- Log + trace + change‑set synthesis in seconds
- Culprit line surfaced, fix drafted mid‑incident
- Triage keeps pace with a moving system
Financial research & security
Conditions don’t wait for throughput. Ultrafast reads market signals, sizes up transactions, and flags suspicious activity while the tape is still moving — complex research that behaves like a real‑time interaction rather than an overnight report.
- Cross‑source synthesis before the bell
- Anomaly triage inside the trading window
- Frontier reasoning at tape speed
Customer support & voice
Complex issues resolved inside the conversation — no dead air, no “let me transfer you.” Even when the answer takes multiple steps across multiple systems, Sol keeps pace with natural speech, so the call never feels like it’s waiting on a model.
- Sub‑second turn‑taking on hard questions
- Multi‑system lookups mid‑sentence
- Voice that sounds fluent because it is
Commerce
A shopper compares two jackets, asks about sizing and stock, hesitates on shipping — and gets every answer before the hesitation hardens into an abandoned cart. Ultrafast answers product questions, checks inventory, personalizes recommendations, and untangles checkout snags while the decision is still live.
- Instant answers at the decision moment
- Inventory + recommendations in one turn
- Checkout issues fixed before abandonment
Live research & experimentation
Work that used to mean launching a batch overnight and reading results with morning coffee compresses into an interactive working session: test an idea, examine results, adjust the approach, run the next experiment — repeatedly, before lunch.
- Overnight batch → same‑day iteration loop
- Search, query, summarize across tools at pace
- The loop tightens around the researcher
05 — CUSTOMER VOICE
What early customers are experiencing.
“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”
John Crepezzi · AI Assistants, Jane Street
“For us the Ultrafast has been invaluable in our voice stack. The speed completely changes the call experience for the more complex work.”
Courtland Lykins · Product Lead — Voice AI, Podium
“Ultrafast allows us to create synchronous experiences for users that were previously limited by intelligence. Oftentimes the barrier to truly fast products is not just tokens per second, but also model intelligence, and ultrafast combines both.”
Mitch Troyanovsky · Co‑Founder, Basis
“Speed doesn’t just make the product feel better. It changes what people can realistically use it for. Ultrafast makes complex financial research feel like a real‑time interaction.”
Alex Wang · Applied AI, Rogo
During the preview we’re working with an initial group of customers across coding, commerce, financial research, support, and interactive applications — studying where an order‑of‑magnitude change in speed creates the most value.
06 — DOGFOOD
How OpenAI is using Ultrafast.
Inside OpenAI, developers have been testing GPT‑5.6 Sol on Ultrafast to find which workflows change most when frontier intelligence answers in real time. Two patterns emerged first.
Incident response, mid‑incident
When an alert fires, engineers need an accurate picture while the system — and the evidence — is still changing. Teams use Ultrafast to read logs, analyze traces, synthesize conversations, line up the next checks, and prepare or validate a fix in a fraction of the usual time. The delay between observing a signal, testing a hypothesis, and choosing the next action shrinks — while engineers remain responsible for judgment and deployment.
Research loops, tightened to the workday
In research, teams use Ultrafast to search knowledge sources, query data, and gather, organize, and summarize information across connected tools at speed. The classic rhythm — launch a batch overnight, read results in the morning — tightens into multiple full iterations inside a single workday.
07 — THE ENGINE
Powered by Cerebras.
Ultrafast marks the next step in OpenAI’s partnership with Cerebras to bring ultra‑low‑latency inference to the platform. Wafer‑scale compute keeps OpenAI’s most intelligent model running at up to 750 output tokens per second — so businesses can build more responsive products, make faster decisions, and put powerful AI directly into their most demanding workflows.
- 750 output tokens / second, up to
- Ultra‑low‑latency inference hardware
- GPT‑5.6 Sol — the full frontier model
08 — AVAILABILITY
In limited preview today. Expanding as capacity grows.
GPT‑5.6 Sol on Ultrafast mode is available now to a select group of customers. If your business needs frontier intelligence at the highest speed, join the list — we notify as access expands.
No spam — one email when your cohort opens up.