AI UX PlaygroundNewsletterJoin 2K+ AI designers and PMs on Substack. New teardowns, patterns, and prompts as they drop.

Inputs

Voice Input

Capture speech, show listening state, and turn audio into text or commands with visible feedback. People need to know when the mic is on, what was heard, and how to correct it.

Interactive demo

Overview

How might we design voice input so people can trust and act on AI output?

When to use

  • Essential for mobile applications, accessibility tools, and hands-free interaction scenarios where voice input provides a more natural and convenient interaction method.

When to skip

  • Silent environments or contexts where speaking is socially impossible.
  • Highly confidential spaces where audio capture is prohibited.
  • Precision editing tasks better served by keyboard alone.

Rules

  • Mic on with no visual or auditory indicator.

  • Finalizing transcript with no chance to edit before send.

  • Auto-sending voice messages without confirmation.

  • No fallback to typing when recognition fails.

Evidence

ProductImplementation
ChatGPTVoice mode with listening UI and transcript.
SiriSystem voice input with waveform and confirmation.
Google AssistantSpeech-to-text with on-screen recognition feedback.
ClaudeVoice input into the chat composer.

Real-world examples

See all

FAQ

What makes good voice input UX?

Clear listening state, live or near-live transcript, easy cancel, and edit-before-send for anything consequential.

How does voice input differ from voice-to-action?

Voice input primarily produces text or a prompt. Voice-to-action maps speech to a specific command or workflow.

Should partial transcripts stream?

Yes when latency allows. Streaming text reassures users they are being heard and lets them stop early.

What about noisy environments?

Show confidence, allow typed fallback, and avoid auto-send when recognition is weak.