Quality and measurement3 min read

Naturalness (MOS)

Also known as: mean opinion score

In short

Naturalness is how human a voice sounds, often summarized as a mean opinion score (MOS): listeners rate samples, and the ratings are averaged. A no-reference model can estimate it automatically. It is separate from accuracy and from similarity.

Why it matters

Naturalness captures the uncanny feeling that the other metrics miss. A voice can say every word correctly and sound exactly like the target speaker and still feel subtly wrong, and that wrongness is what a naturalness score is trying to put a number on.

Because it started as a human panel score, it is inherently subjective, but averaging across many listeners makes it stable enough to compare systems, and automated estimators now let you screen without a panel every time.

In practice

Use naturalness as one axis among several, never alone. A high MOS with a poor word error rate means a fluent voice that gets words wrong. Screen with an automated estimator, then trust your own ears on the passages that matter.

Related terms

Related reading

Blog
Natural enough to publish

Your voice, on your Mac

Vocast clones your voice from about ninety seconds and narrates any script in it, fully on-device, for $49 one time.