Cloning and synthesis2 min read

Text-to-speech (TTS)

Also known as: TTS

In short

Text-to-speech turns written text into spoken audio using a synthetic voice. Modern TTS is almost always intelligible; the interesting differences are in prosody and in whether the voice is a generic one or a clone of a specific speaker.

Why it matters

Text-to-speech is the plumbing under a lot of everyday audio: screen readers, assistants, announcements, and now long-form narration. Because it is so widely used, the bar has moved from can you understand it, which almost everything now clears, to does it sound like a person you would listen to for twenty minutes.

That is why two TTS voices with identical clarity can feel completely different: the gap is prosody and identity, not intelligibility.

In practice

Judge a TTS voice on a real paragraph, not a polished demo sentence. Feed it numbers, names, and loanwords, which is where word errors hide, and listen to a full passage for flatness and pacing rather than a single clean line.

Related terms

Your voice, on your Mac

Vocast clones your voice from about ninety seconds and narrates any script in it, fully on-device, for $49 one time.