Live Transcript

Live transcript is an AI UX pattern that shows speech-to-text in near real time during voice or meeting capture. Users can verify recognition, quote accurately, and resume by reading—not only by listening.

Share

Interactive demo

00:12...andthat'swhyweprioritizeuserexperienceaboveallelse

Overview

The design problem

How might we design live transcript so people can trust and act on AI output?

Use this pattern

When this pattern fits

  • Perfect for meeting tools, accessibility applications, and voice-enabled platforms where real-time transcription makes audio content accessible and searchable.

Avoid this pattern

When to skip or lighten it

  • Silent text-only products.
  • Ultra-low-latency voice UIs where transcript lag would confuse more than help.
  • Highly confidential calls where storing transcript text is forbidden.

States

State model coming soon

Key UX elements

Key UX elements coming soon

Anti-patterns to avoid

  • Final-only transcripts with no live feedback while speaking.

  • Transcripts that cannot be edited after recognition errors.

  • No speaker labels in multi-person calls.

  • Auto-saving transcripts without consent.

How products use it

ProductImplementation
Otter.aiLive meeting transcripts with speakers.
ZoomLive transcription during calls.
Microsoft TeamsReal-time captions and transcripts.
Google MeetCaptions and transcript affordances.

Implementation

Copy this prompt to generate a production-ready implementation in Cursor, Claude Code, Lovable, or any AI coding agent.

Generate a production-ready implementation of the "Live Transcript" AI interface design pattern.

Pattern Definition:

Frequently asked questions

Why show a live transcript in voice AI?

So users can catch mishears early, interrupt, or edit before the system acts on wrong words.

Partial vs final transcripts?

Show partials for responsiveness, then stabilize finals. Style partials differently so users know they may change.

How does this relate to voice input?

Voice input is the capture mode. Live transcript is the readable feedback channel for that capture.

Multi-speaker necessities?

Speaker labels and diarization when more than one person talks; otherwise attribution breaks.

Weekly AI UX in your inbox

Weekly AI interface UX notes and resources on Substack, no spam, unsubscribe anytime.

Subscribe on Substack