Naturalness (MOS)
Also known as: mean opinion score
Naturalness is how human a voice sounds, often summarized as a mean opinion score (MOS): listeners rate samples, and the ratings are averaged. A no-reference model can estimate it automatically. It is separate from accuracy and from similarity.
Why it matters
Naturalness captures the uncanny feeling that the other metrics miss. A voice can say every word correctly and sound exactly like the target speaker and still feel subtly wrong, and that wrongness is what a naturalness score is trying to put a number on.
Because it started as a human panel score, it is inherently subjective, but averaging across many listeners makes it stable enough to compare systems, and automated estimators now let you screen without a panel every time.
In practice
Use naturalness as one axis among several, never alone. A high MOS with a poor word error rate means a fluent voice that gets words wrong. Screen with an automated estimator, then trust your own ears on the passages that matter.