AI UX PlaygroundNewsletterJoin 2K+ AI designers and PMs on Substack. New teardowns, patterns, and prompts as they drop.

Audio

Live Transcript

Show speech-to-text as people talk or record a meeting. They can verify recognition, quote accurately, and catch up by reading instead of replaying audio.

Interactive demo

Live transcript

00:12...and that's why we prioritize user experience above all else

Overview

How might we design live transcript so people can trust and act on AI output?

When to use

  • Perfect for meeting tools, accessibility applications, and voice-enabled platforms where real-time transcription makes audio content accessible and searchable.

When to skip

  • Silent text-only products.
  • Ultra-low-latency voice UIs where transcript lag would confuse more than help.
  • Highly confidential calls where storing transcript text is forbidden.

Rules

  • Final-only transcripts with no live feedback while speaking.

  • Transcripts that cannot be edited after recognition errors.

  • No speaker labels in multi-person calls.

  • Auto-saving transcripts without consent.

Evidence

ProductImplementation
Otter.aiLive meeting transcripts with speakers.
ZoomLive transcription during calls.
Microsoft TeamsReal-time captions and transcripts.
Google MeetCaptions and transcript affordances.

FAQ

Why show a live transcript in voice AI?

So users can catch mishears early, interrupt, or edit before the system acts on wrong words.

Partial vs final transcripts?

Show partials for responsiveness, then stabilize finals. Style partials differently so users know they may change.

How does this relate to voice input?

Voice input is the capture mode. Live transcript is the readable feedback channel for that capture.

Multi-speaker necessities?

Speaker labels and diarization when more than one person talks; otherwise attribution breaks.