I thought multi-speaker TTS would be simple

by

I thought multi-speaker TTS would be simple.

Give each character a voice.
Paste the dialogue.
Generate the audio.

But while building TTS Dialog, I realized the annoying part isn’t generating speech.

It’s everything around it.

Changing speakers line by line.
Regenerating one bad sentence without touching the rest.
Keeping voices consistent across a long conversation.
Figuring out which voice actually fits each character.

So I’m curious:

If you create content with multiple voices — lessons, podcasts, stories, videos, games, anything —

what’s the most annoying part of your current workflow?

I’m building around this problem right now, so I’d genuinely love to steal… I mean, learn from your workflow 😅

If anyone wants to play with what I’ve built so far:

3 views

Add a comment

Replies

Be the first to comment