New speed class · Now in limited preview

Ultrafast mode. GPT‑5.6 Sol at up to 14× the speed.

Frontier intelligence at 750 output tokens per second — the same GPT‑5.6 Sol, up to fourteen times faster than Standard processing, launching first in the OpenAI API. Powered by Cerebras.

0 tok/s · peak
0.0× vs Standard
750 output tokens/s · up to
1 model · no downgrade
SCROLL

01 — THE TRADE‑OFF IS OVER

Real‑time speed used to mean a smaller, dumber model.

Until now, getting real‑time speed typically meant choosing a smaller or more specialized model — trading away the intelligence your hardest problems need. Ultrafast points in a new direction: more useful work per second.

When speed no longer requires giving up intelligence, AI moves into the most time‑sensitive parts of a business — and new kinds of work become possible. The same GPT‑5.6 Sol frontier model, running up to 14× faster, on ultra‑low‑latency hardware.

Specifications

0 tok/s

Peak output throughput — frontier intelligence at reading‑defying speed.

0× faster

Versus Standard processing on GPT‑5.6 Sol — roughly 54 tokens per second.

GPT‑5.6 Sol

Our most intelligent model — unchanged. No smaller stand‑in, no distilled substitute.

API first

Launching first in the OpenAI API, in limited preview for select customers.

02 — THE CENTERPIECE

Same prompt. Same brain. Different clock.

Watch GPT‑5.6 Sol answer the exact same request in Ultrafast and Standard modes, token by token, at published rates. This is what fourteen times looks like.

ULTRAFAST · GPT‑5.6 SOL 0 tok/s READY

          

Ultrafast lane idle

STANDARD · GPT‑5.6 SOL 0 tok/s READY

          

Standard lane idle

Simulated locally at published rates — up to 750 tok/s vs ~54 tok/s. Answers are illustrative, scripted output.

03 — THE THROTTLE

Push it anywhere from 54 to 750 tok/s.

Pick a task or write your own, drag the throttle, and watch the same intelligence answer at any speed in the new class — with Standard processing along for comparison.

54 750
750 tok/s · full send
0tokens
0.0selapsed
0tok/s live
OUTPUT · ULTRAFAST IDLE
Pick a task and hit Generate — the answer streams at the throttle speed.

Playground idle

ULTRAFAST 750 tok/s
STANDARD ~54 tok/s

Demo simulation — custom prompts get a scripted illustrative answer, streamed character‑for‑token at your chosen rate.

04 — WHERE SPEED WINS

The most time‑sensitive work in your business, unblocked.

Five workflows where an order‑of‑magnitude jump in speed changes what the product even is.

01

Incident response & reliability

When a critical system fails, Sol reads the application logs, the deploy diff, and the on‑call thread while the outage is still unfolding — and hands back a ranked hypothesis with a candidate fix before the first bridge call warms up. Mean time to understand collapses; engineers stay in charge of judgment and deployment.

  • Log + trace + change‑set synthesis in seconds
  • Culprit line surfaced, fix drafted mid‑incident
  • Triage keeps pace with a moving system
02

Financial research & security

Conditions don’t wait for throughput. Ultrafast reads market signals, sizes up transactions, and flags suspicious activity while the tape is still moving — complex research that behaves like a real‑time interaction rather than an overnight report.

  • Cross‑source synthesis before the bell
  • Anomaly triage inside the trading window
  • Frontier reasoning at tape speed
03

Customer support & voice

Complex issues resolved inside the conversation — no dead air, no “let me transfer you.” Even when the answer takes multiple steps across multiple systems, Sol keeps pace with natural speech, so the call never feels like it’s waiting on a model.

  • Sub‑second turn‑taking on hard questions
  • Multi‑system lookups mid‑sentence
  • Voice that sounds fluent because it is
04

Commerce

A shopper compares two jackets, asks about sizing and stock, hesitates on shipping — and gets every answer before the hesitation hardens into an abandoned cart. Ultrafast answers product questions, checks inventory, personalizes recommendations, and untangles checkout snags while the decision is still live.

  • Instant answers at the decision moment
  • Inventory + recommendations in one turn
  • Checkout issues fixed before abandonment
05

Live research & experimentation

Work that used to mean launching a batch overnight and reading results with morning coffee compresses into an interactive working session: test an idea, examine results, adjust the approach, run the next experiment — repeatedly, before lunch.

  • Overnight batch → same‑day iteration loop
  • Search, query, summarize across tools at pace
  • The loop tightens around the researcher

05 — CUSTOMER VOICE

What early customers are experiencing.

“The increase in speed brought by Cerebras is impressive. It enables different ways of using the models, and makes it practical for developers to work in a more focused and productive way alongside them.”

John Crepezzi · AI Assistants, Jane Street

During the preview we’re working with an initial group of customers across coding, commerce, financial research, support, and interactive applications — studying where an order‑of‑magnitude change in speed creates the most value.

06 — DOGFOOD

How OpenAI is using Ultrafast.

Inside OpenAI, developers have been testing GPT‑5.6 Sol on Ultrafast to find which workflows change most when frontier intelligence answers in real time. Two patterns emerged first.

Incident response, mid‑incident

When an alert fires, engineers need an accurate picture while the system — and the evidence — is still changing. Teams use Ultrafast to read logs, analyze traces, synthesize conversations, line up the next checks, and prepare or validate a fix in a fraction of the usual time. The delay between observing a signal, testing a hypothesis, and choosing the next action shrinks — while engineers remain responsible for judgment and deployment.

Research loops, tightened to the workday

In research, teams use Ultrafast to search knowledge sources, query data, and gather, organize, and summarize information across connected tools at speed. The classic rhythm — launch a batch overnight, read results in the morning — tightens into multiple full iterations inside a single workday.

07 — THE ENGINE

Powered by Cerebras.

Ultrafast marks the next step in OpenAI’s partnership with Cerebras to bring ultra‑low‑latency inference to the platform. Wafer‑scale compute keeps OpenAI’s most intelligent model running at up to 750 output tokens per second — so businesses can build more responsive products, make faster decisions, and put powerful AI directly into their most demanding workflows.

  • 750 output tokens / second, up to
  • Ultra‑low‑latency inference hardware
  • GPT‑5.6 Sol — the full frontier model

08 — AVAILABILITY

In limited preview today. Expanding as capacity grows.

GPT‑5.6 Sol on Ultrafast mode is available now to a select group of customers. If your business needs frontier intelligence at the highest speed, join the list — we notify as access expands.

No spam — one email when your cohort opens up.

View more demos Get $10 off Kimi K3