Overview
How might we make generation feel immediate without forcing people to wait for a finished answer?
When to use
- Chat, coding, and writing tools where responses take seconds and blank waits feel broken.
- Long-form answers where people skim early and may change direction mid-stream.
- Any surface where time-to-first-token matters more than total generation time.
- Products that need interruptibility: stop, edit, or retry before a full wrong answer finishes.
When to skip
- Tiny deterministic outputs (a single number or yes/no) where streaming adds noise.
- Layouts that must paint as a complete unit (some charts or print-ready pages) until structure is known.
- Accessibility contexts where rapid token updates overwhelm screen readers without a “complete” announcement.
States
Design the whole generation lifecycle, not only the blinking caret.
- 01
Waiting
The request is in flight but no tokens yet. A thinking status proves work started without faking the answer.
- 02
Streaming
Tokens append in place. Layout stays stable; a caret or live cue marks the growing end.
- 03
Interrupted
Stop cancels further tokens. Partial text remains visible and usable.
- 04
Complete
Generation finished. The caret clears; copy, share, and follow-ups unlock on the final text.
- 05
Failed
The stream errored. Offer retry without losing the prompt or prior turns.
Key UX elements
The parts that must be present for streaming to feel fast and controllable.
First token
Prove work started quickly.
Show a thinking status before tokens arrive. Empty silence after send feels like failure.
Caret
Mark the live edge of the reply.
A subtle caret or pulse shows growth without competing with the words.
Stop
Let people cancel mid-answer.
A visible Stop control is table stakes when the direction is wrong before the end.
Layout
Keep the page from thrashing.
Stream into a stable block. Jumping scroll and reflow on every token breaks reading.
Partial
Treat early text as useful.
Allow skim, copy, and redirect before the full answer finishes when the draft is already clear.
Complete
Announce the end, not every token.
Clear the live cue and prefer a single “response ready” signal for assistive tech.
Rules
No stop or cancel control while tokens stream.
Layout thrash that jumps the page with every token.
Streaming fake progress while the model has not started.
Blocking copy/share until the full answer finishes when partial text is already useful.
Evidence
| Product | Implementation |
|---|---|
| ChatGPT | Token stream in the thread with stop generation during the run. |
| Claude | Streaming replies with interrupt and retry on the same turn. |
| Cursor | Streams code and chat; apply/diff flows wait for stable hunks. |
| Perplexity | Streams the answer while research steps and citations settle. |
