A voice model can score well and still feel terrible to use. Why?
I work on both ASR and TTS evaluation, and one thing I’ve learned the hard way is that model quality and product quality are not the same thing.
You can have good recognition accuracy and still get an output that makes the user think: “I would’ve been faster typing this.”
The failures I notice most aren’t always obvious transcription errors.
Sometimes every word is technically correct, but the system misses the intent.
Sometimes it preserves every filler word and self-correction so faithfully that the output is unusable.
Sometimes a proper noun is wrong and that one error matters more than ten minor mistakes.
And in multilingual speech, a model that looks strong overall can behave completely differently once users start switching languages naturally.
This has made me question how much traditional speech metrics really tell us about product experience.
For people building voice products: what do you measure beyond raw recognition accuracy?
I’m especially curious whether anyone has found a good way to quantify “repair cost” — how much work the user has to do after the model produces its answer.

Replies