AI UX PlaygroundNewsletterJoin 2K+ AI designers and PMs on Substack. New teardowns, patterns, and prompts as they drop.

Audio

Voice Cloning

Build a custom TTS voice from user samples, with training progress, quality checks, and preview before use. Keep consent and rights clear, and keep humans in control of deployment.

Interactive demo

Voice Cloning
Record Sample Voice
Record at least 10 seconds of speech

Overview

How might we design voice cloning so people can trust and act on AI output?

When to use

  • Ideal for content creation tools, accessibility applications, and personalized voice assistants where custom voice generation enhances user experience.

When to skip

  • No verified consent from the voice owner.
  • Regions restrict synthetic voice without disclosure.
  • Sample audio is too noisy for acceptable clone quality.

Rules

  • Clone from a few seconds with no quality warning.

  • No watermark or label on synthetic speech output.

  • Celebrity or third-party voices without rights checks.

  • Training UI that hides how samples will be stored.

Evidence

ProductImplementation
ElevenLabsInstant voice clone from uploads with stability sliders.
DescriptOverdub voice model trained on user recordings.
Resemble.aiCustom voice APIs with consent and verification flows.
Play.htVoice library plus clone-from-sample for creators.

FAQ

How much audio is needed to clone a voice?

Products range from ~30 seconds for rough clones to several minutes for stable commercial quality. Show requirements upfront.

What consent UX is required?

Recorded attestation, terms acceptance, and optional liveness check before training starts.

How label cloned output?

Disclose synthetic speech in UI and metadata, especially for public or phone-facing use.

Clone vs preset voices?

Presets are licensed stock. Clones are user-specific models trained on their samples.