I thought multi-speaker TTS would be simple
I thought multi-speaker TTS would be simple.
Give each character a voice.
Paste the dialogue.
Generate the audio.
But while building TTS Dialog, I realized the annoying part isn’t generating speech.
It’s everything around it.
Changing speakers line by line.
Regenerating one bad sentence without touching the rest.
Keeping voices consistent across a long conversation.
Figuring out which voice actually fits each character.
So I’m curious:
If you create content with multiple voices — lessons, podcasts, stories, videos, games, anything —
what’s the most annoying part of your current workflow?
I’m building around this problem right now, so I’d genuinely love to steal… I mean, learn from your workflow 😅
If anyone wants to play with what I’ve built so far:
https://app.ttsdialog.com
Replies