AI UX PlaygroundNewsletterJoin 2K+ AI designers and PMs on Substack. New teardowns, patterns, and prompts as they drop.

Deep dives · 9 min

The Eight States Every AI Feature Needs Before Launch

Empty, thinking, streaming, partial, error, refusal, handoff, and limits: what each state must tell people, shipped examples, and how to test it before launch.

Key takeaways

  • An AI feature is not one screen. Between open and result it passes through empty, thinking, streaming, partial, error, refusal, handoff, and limit states.
  • Each state answers a question the person is already asking, such as "Is it working?" or "Is my work safe?" Silence makes them guess.
  • Every run should end in a result, a retry, or an explanation. A thinking state with no exit is the failure to design against.
  • Partial results are the most common outcome of multi-step work. Lead with counts and retry only what failed.
  • A state passes launch review when someone can trigger it on demand and the result meets the bar.

Demos show the happy path. Real use happens in the states around it: a blank box, a long wait, half an answer, a refusal, a limit. Model calls take seconds, not milliseconds. Output arrives in pieces. Requests fail, get refused, or run out of quota.

When the interface stays silent in those states, people guess. They resend the request, abandon the task, or stop trusting results that were fine. Each section below covers one state: the question it answers, what to design, a shipped example, what to avoid, and how to trigger it in QA. A launch review and a copyable checklist close the guide.

1. Empty

What can this do for me?

Appears on first use, in a new conversation or blank document, and after someone clears the input.

The empty state is the feature's only chance to set expectations before the first request. A blank box that says "Ask me anything" invites requests the feature can't handle, and the first failure costs more trust than any later one.

Design for it

  • State what the feature is for in one line, in the person's terms rather than the model's.
  • Offer three to six starter tasks built from the person's own context, such as the open file, a recent meeting, or their role.
  • Show the input the feature expects: a question, a file, a URL, or a selection.
  • Surface any limit that would break a first try, like file size or supported languages.
  • Put the cursor in the input and show the keyboard shortcut.

Shipped example: Lovable

Lovable's home screen is almost entirely its prompt box. The placeholder types out a complete request ("Ask Lovable to build…"), and the attach menu and the Build/Plan switch sit inside the box. The first action and its options live on one surface, and the placeholder teaches the shape of a good request without a tutorial.

Real-world example

Lovable home screen with a personalized greeting and a prompt box containing the attach menu, Build mode, mic, and send

Lovable · One surface for the first action. The placeholder models a real request, and templates sit below. Full teardown

Avoid

  • An open-ended prompt with no hint of scope.
  • Starter tasks that can't succeed with this person's data or plan.
  • Onboarding carousels that delay the first real request.

Test it: Open the feature from a brand-new account with no data. Someone should complete one useful task in under two minutes without reading the docs.

Pattern: Prompt Startersstarter tasks that teach the shape of a good request.

2. Thinking

Is it working, and on what?

Appears between submit and the first visible output. For reasoning models and agents, it can last from a second to several minutes.

Silence reads as failure. After a few seconds without feedback, people click again, refresh, or leave, and a duplicate request can double the cost or the side effects.

Design for it

  • Acknowledge the request immediately: show the message as sent and block duplicate submits.
  • Within a second or two, replace a generic spinner with plain-language steps, such as "Searching 12 sources" or "Reading contract.pdf".
  • For waits over about ten seconds, show elapsed time and offer to notify the person when it's done.
  • Keep Stop visible the whole time.
  • Make sure every run ends in a result, a retry, or an explanation.

Shipped example: Perplexity

While it answers, Perplexity lists the steps it takes. Once the answer arrives, those steps collapse into a short "Completed 2 steps" summary above the answer. The process stays one click away without crowding the result. Nielsen Norman Group's Perplexity interview covers the same design choice.

Real-world example

Perplexity answer with a collapsed Completed 2 steps summary showing Searching the web above the response

Perplexity · Steps collapse into a one-line summary once the answer lands. The process stays one click away. Full teardown

For multi-minute agent runs, Manus sets the expectation in words ("It may take a few minutes"), names the current step, and shows an elapsed timer beside the composer.

Real-world example

Manus working state with an Editing files step card and an elapsed timer reading 0:08

Manus · A named step, an honest time estimate, and elapsed time for a long run. Full teardown

The failure to design against is a thinking state with no exit. ChatGPT users have reported long reasoning runs that end with "Stopped thinking" and no saved response. Whatever interrupts the work, the person should get something they can act on.

Avoid

  • An indeterminate spinner for more than a few seconds.
  • Progress bars that don't track real work.
  • Raw reasoning or tool logs as the default view.

Test it: Replay requests at your p95 and p99 latency. Take a screenshot at 1, 10, and 60 seconds. Each one should say what is happening and offer Stop.

Pattern: Progress Stepsplain-language steps that collapse once the answer arrives.

3. Streaming

Can I read it yet, and can I stop it?

Appears while output arrives in pieces, token by token or step by step.

Streaming makes waits feel shorter, but it creates its own problems: layouts that jump as tables and code blocks render, auto-scroll that fights someone who is reading, and an end state that isn't obvious.

Design for it

  • Turn Send into Stop while output streams. Stopping keeps what has already arrived.
  • If the person scrolls up, stop auto-scrolling until they return to the bottom.
  • Reserve space for elements that arrive late, such as citations, code blocks, and tables.
  • Make done unambiguous: Stop turns back into Send, and actions like Copy and Retry appear.
  • For screen readers, announce when the response starts and when it finishes, not every token.

Shipped example: ChatGPT and Claude

Both products turn the send button into a stop button while a response streams, so the control that started the request is the one that ends it. Message actions such as copy and retry belong to the finished response.

Avoid

  • Copy or Share available mid-stream, so people copy half an answer.
  • Markdown that reflows as it renders, such as tables snapping from pipes into a grid.
  • Auto-scroll that pulls the reader back down.

Test it: Ask for a long answer with a table and a code block. Scroll up mid-stream, press Stop halfway through, then repeat with VoiceOver or NVDA running.

Pattern: Streaminginterruptible output with a clear end state.

4. Partial

What got done, and what didn't?

Appears when an agent finishes some steps but not others, a batch job processes most items, or an answer covers part of the question.

Partial results are the most common outcome of multi-step work and the least designed. A success message over a half-finished job is worse than an error, because nobody goes back to check.

Design for it

  • Lead with counts: "Updated 42 of 48 listings."
  • List what didn't happen, item by item, with the reason.
  • Keep completed work, and retry only the failed part.
  • Say whether anything is left in an inconsistent state, and how to fix it.
  • Leave a place to resume: a link, a task, or a saved state.

Shipped example: GitHub Copilot coding agent

When someone assigns a task to GitHub's Copilot coding agent, it makes the changes in a pull request and keeps updating the description as it works, so reviewers can see what's done and what remains before the task finishes. The pull request doubles as the resume point: people review, comment, and ask for changes in the same place.

Real-world example

Copilot's pull request description with a seven-item task checklist, from analyzing the codebase to reviewing the implementation

GitHub Copilot · The plan lives in the pull request as a checklist Copilot ticks off as it works, so what's done and what's left stay visible. GitHub Skills

Avoid

  • A generic "Done" when some items failed.
  • A retry that reruns completed steps and repeats side effects, such as sending an email twice.
  • Failures that only show up in a log.

Test it: Run a batch where a known subset fails: invalid rows, a missing permission, a timeout on item 30. The summary should name each failure, and retry should skip what already succeeded.

Pattern: Tool Useper-item status so failures stay visible and retryable.

5. Error

What broke, and is my work safe?

Appears when the model call fails, a tool returns an error, the output is invalid, or the connection drops.

AI errors are harder to explain than ordinary ones, because the cause is often unclear to the product too. People need the cause less than they need two answers: is my work safe, and what should I do now?

Design for it

  • Say what happened in plain language, and whether retrying is likely to help.
  • Confirm what's preserved: the input, earlier output, and any files.
  • Offer one primary recovery action: retry, fix, or restore.
  • When the AI changed something, give people a restore point.
  • Keep error codes and logs available for support, but secondary.

Shipped examples: Lovable and Cursor

Two shipped approaches cover most cases. Lovable fixes forward: when a build error appears, a Try to fix button has the agent read the logs and attempt a repair. Cursor rolls back: as the agent edits, it creates checkpoints, and Restore Checkpoint reverts the files changed after that point while keeping the conversation, so nothing the person said is lost.

Real-world example

Cursor chat panel with a Restore checkpoint button under an earlier message

Cursor · Restore checkpoint sits on the earlier message, so rolling back the agent's file changes is one click from where they started. Cursor Forum

Avoid

  • "Something went wrong" with no next step.
  • Clearing the input when a request fails.
  • Silent automatic retries that repeat work or spend credits.

Test it: Cut the network mid-request, return a malformed tool response, and force a server error from the model provider. After each one, the input should still be there and exactly one next action should be obvious.

Pattern: Checkpoints and Restorea restore point for anything the AI changed.

6. Refusal

Why not, and what can I do instead?

Appears when a request conflicts with policy, goes beyond what the feature can do, or needs a permission the person doesn't have.

Refusals feel personal in a way errors don't. A lecture, or a refusal of the whole request when only one part was a problem, teaches people either to rephrase until something slips through or to stop asking.

Design for it

  • Say no once, briefly, without moralizing.
  • Name the kind of no: won't (policy), can't (capability), or not allowed (permission). Each needs a different next step.
  • Offer the closest thing you can do.
  • Still help with the parts of the request that are fine.
  • Give people a way to report a refusal that seems wrong.

Shipped example: OpenAI

OpenAI has moved its models away from flat refusals. With GPT-5 it introduced safe completions, which aim for the most helpful response within safety limits instead of a hard no, and it has said it is working to reduce unnecessary refusals and preachy responses in ChatGPT. The lesson for product teams: the best refusal is often a partial yes.

You can see the capability version in ChatGPT itself. Asked to book a restaurant, it says it can't make reservations in the first sentence, then offers nearby options and asks the questions it needs to help with the rest.

Real-world example

ChatGPT says it can't make reservations, suggests two nearby restaurants, and asks about party size, time, cuisine, and seating

ChatGPT · A capability no in one line, then the closest yes: suggestions and the questions needed to finish the job. Full teardown

Avoid

  • Paragraphs of policy explanation.
  • A watered-down answer that doesn't say it was watered down.
  • The same message for policy, capability, and permission limits.

Test it: Keep a set of prompts that are clearly fine, clearly not, and borderline. Track how often the clearly fine ones get refused, and check that every refusal offers a next step.

7. Handoff

Who's helping me now, and do they know what I said?

Appears when the AI passes the task to a person, another team, or another tool.

A good handoff feels like a colleague walking your question over. A bad one feels like starting again, and it erases the time the AI saved.

Design for it

  • Say who takes over and when to expect a reply, such as "A billing specialist, usually within two hours."
  • Carry the context: the conversation, what was tried, and account details.
  • Let people ask for a person directly, and honor it the first time.
  • Show status after the handoff: queued, assigned, replied.
  • Set expectations outside working hours.

Shipped example: Intercom Fin

Intercom lets support teams decide when Fin escalates, whether it offers escalation or escalates immediately, and what Fin says during the handover. Fin and human teammates work from the same conversation record, so whoever picks up the conversation sees what the customer already said.

Real-world example

Intercom workflow where Fin's handover branches by plan: free users get a community link, everyone else sees 'An agent will be with you soon' and is assigned to Tier 2 Support

Intercom Fin · When Fin hands over, the workflow says who's next and routes the conversation: a human for most plans, the community for free users. Intercom Help

Avoid

  • Handing over without the transcript.
  • Asking the same questions again after the handoff.
  • A "Talk to a person" option that loops back to the bot.

Test it: Trigger every escalation rule. Then act as the receiving teammate: you should be able to resolve the issue without asking the customer anything they already said.

Pattern: Human Handofftransfer the job with its context when the AI can't continue.

8. Limits

When can I continue?

Appears when someone reaches a usage cap, rate limit, context window, file-size limit, or plan quota.

Limits are business decisions, but people meet them as interface states, often in the middle of a task. What the interface says is most of the difference between a limit that feels fair and one that feels hostile.

Design for it

  • Warn before the wall, while there's still room to finish the current task.
  • Say exactly when the limit resets, in the person's local time.
  • Say what still works: other models, history, drafts, or exports.
  • Preserve work in progress.
  • Offer options (wait, switch, or upgrade), and treat waiting as a real option.

Shipped example: Claude

Claude's session limits reset every five hours. When someone reaches one, the limit message states when it resets. Paid members with usage credits enabled see a variant that says Claude is continuing on credits, so reaching the limit changes billing instead of stopping the work.

Warning before the wall can be as simple as Manus's credits popover: the current balance, the daily refresh amount, and the exact refresh time, one click from the composer.

Real-world example

Manus credits popover showing 1,000 free credits and a daily refresh to 300 at 00:00

Manus · Balance, refresh amount, and refresh time in one place, before anyone hits the limit. Full teardown

Avoid

  • A hard stop mid-task with no warning.
  • "Try again later" without a time.
  • Losing an in-progress response when the limit hits.

Test it: Give a test account a tiny quota. Hit the limit mid-response and mid-agent-run. The person should see the reset time and keep everything they had.

Pattern: Rate Limit Warningswarn before the wall and say when it resets.

Launch review

Run this review with design, engineering, and support in the room. A state passes when someone can trigger it on demand and the result meets the bar.

StateTriggerPasses when
EmptyA new account with no dataSomeone completes a useful task in under two minutes without the docs
ThinkingRequests replayed at p95 and p99 latencyThe screen explains the wait at 1, 10, and 60 seconds and offers Stop
StreamingA long answer with a table and codeStop keeps output, scrolling up pauses auto-scroll, and screen readers hear the end
PartialA batch with known failuresThe summary counts successes, names each failure, and retry skips finished items
ErrorNetwork cut, malformed tool output, provider errorInput is preserved, one recovery is obvious, and AI changes can be restored
RefusalClearly fine, clearly not, and borderline promptsRefusals are brief, name the reason type, and offer an alternative
HandoffEvery escalation ruleThe receiver has full context, and the person can see status
LimitsA test account with a tiny quotaPeople are warned first, see the reset time, and keep their work

Checklist

Copy it into a launch doc, pull request template, or tracker. The Markdown checkboxes work in GitHub, Linear, and Notion.

AI feature launch checklist
## AI feature launch checklist: eight states

### 1. Empty
- [ ] One-line purpose, in the person's terms
- [ ] 3–6 starter tasks built from the person's own context
- [ ] The expected input is visible (question, file, URL, selection)
- [ ] Limits that would break a first try are stated up front

### 2. Thinking
- [ ] Request acknowledged immediately; duplicate submits blocked
- [ ] Plain-language steps within 1–2 seconds
- [ ] Elapsed time and a notify option after about 10 seconds
- [ ] Stop is always visible
- [ ] Every run ends in a result, a retry, or an explanation

### 3. Streaming
- [ ] Send becomes Stop; stopping keeps the partial output
- [ ] No auto-scroll after the person scrolls up
- [ ] Space reserved for late elements (citations, code, tables)
- [ ] Done is unambiguous; message actions appear when complete
- [ ] Screen readers hear the start and the end, not every token

### 4. Partial
- [ ] The outcome leads with counts ("42 of 48")
- [ ] Each failed item is listed with a reason
- [ ] Completed work is kept; retry runs only the failed part
- [ ] Any inconsistent state is named, with the fix
- [ ] There is a place to resume

### 5. Error
- [ ] Plain-language cause, and whether retrying will help
- [ ] Input, earlier output, and files confirmed safe
- [ ] One primary recovery: retry, fix, or restore
- [ ] A restore point for anything the AI changed
- [ ] No silent retries that repeat work or spend credits

### 6. Refusal
- [ ] Brief, without moralizing
- [ ] The kind of no is clear: policy, capability, or permission
- [ ] The closest possible alternative is offered
- [ ] The acceptable parts of the request still get done
- [ ] A way to report a wrong refusal

### 7. Handoff
- [ ] Who takes over, and when to expect a reply
- [ ] Full context travels with the handoff
- [ ] Asking for a person works the first time
- [ ] Status is visible after the handoff
- [ ] Out-of-hours expectations are set

### 8. Limits
- [ ] A warning before the limit, with room to finish
- [ ] The reset time, in the person's local time
- [ ] What still works is stated
- [ ] Work in progress is preserved
- [ ] Waiting is a real option, not only upgrading

Sources

Explore related reference