Voice Visualizer

Voice visualizer is an AI UX pattern that shows real-time motion or waveform feedback during listen, think, and speak states in voice mode. It replaces silent gaps with clear state cues so users know when to talk and when the assistant is responding.

Share

Interactive demo

Overview

The design problem

How might we design voice visualizer so people can trust and act on AI output?

Use this pattern

When this pattern fits

  • Essential for voice assistants, voice-first applications, and hands-free interfaces where visual feedback enhances user understanding of voice interaction states.

Avoid this pattern

When to skip or lighten it

  • Pure text chat with no voice path.
  • Accessibility settings request reduced motion.
  • Background voice where visual feedback is irrelevant.

States

State model coming soon

Key UX elements

Key UX elements coming soon

Anti-patterns to avoid

  • Same animation for listening and speaking, so state is ambiguous.

  • Visualizer with no caption or icon for deaf or low-vision users.

  • Flashy motion that distracts from transcript content.

  • No idle state when mic is off but UI looks “live”.

How products use it

ProductImplementation
SiriOrb and waveform states on iPhone and HomePod screens.
Google AssistantAnimated dots and color shifts for listen vs respond.
AlexaLight ring and screen pulse tied to audio activity.
ChatGPTVoice mode orb with listening and speaking animations.

Implementation

Copy this prompt to generate a production-ready implementation in Cursor, Claude Code, Lovable, or any AI coding agent.

Generate a production-ready implementation of the "Voice Visualizer" AI interface design pattern.

Pattern Definition:

Frequently asked questions

What states should a voice visualizer show?

At minimum: idle, listening, processing, and speaking. Optional: error, muted, and wake-word armed.

Is a visualizer required if you have a transcript?

Transcripts help content; visualizers help timing. Pair both for voice-first UX.

How do you respect reduced motion?

Offer static icons or subtle opacity changes instead of looping waveforms.

What about latency?

Show processing state within ~200ms of end-of-speech so users do not talk over the model.

Weekly AI UX in your inbox

Weekly AI interface UX notes and resources on Substack, no spam, unsubscribe anytime.

Subscribe on Substack