Voice cloning
Also known as: voice replication, speaker cloning, voice model
Voice cloning is the process of building a reusable model of one person's voice from recordings of them speaking, then using that model to read new text aloud. A good clone keeps the timbre, accent, and rhythm of the original speaker, not only the pitch.
Why it matters
Recording narration by hand is the slowest part of publishing audio, and it does not scale. Every edit to the script means going back to the microphone, matching the old take, and re-recording the room. A voice clone turns narration into a text problem, so a change to a sentence costs a sentence of render time instead of a whole session.
It also matters for the people who cannot record on demand: creators who have lost their voice, teams that publish in several languages, and anyone who ships updates every week and needs a voice that stays available.
In practice
A usable clone starts from a clean, dry sample of your own voice, and a minute or two is often enough. The quality question is not whether it sounds like you for one sentence, but whether it holds up across a whole chapter, where flatness and drift show. Judge a clone on long, expressive material, not a demo line.
How Vocast handles this
Vocast builds a voice profile from about ninety seconds of your own speech and narrates scripts up to 20,000 characters in it, entirely on your Mac, with nothing uploaded. It clones only your own voice or a voice you have consent to use.