VocalVia
Turn documents and articles into editable multi-voice audio
109 followers
Turn documents and articles into editable multi-voice audio
109 followers
VocalVia turns PDFs, Word files, Markdown, web articles, and pasted text into structured outlines, editable podcast scripts, and natural multi-voice audio. Choose speakers and voices, refine individual segments, then export the finished audio. It is built for saved reading, study notes, research papers, and long-form content you want to listen to away from a screen.







VocalVia
this is one of the more thoughtful takes on document-to-audio I've seen, mostly because the editable script step exists at all instead of piping straight to TTS. one thing I didn't see covered yet: pronunciation of jargon, acronyms, and author names in research papers. that's usually where TTS breaks the listening experience even when the sentence structure is fine - a mispronounced term every few minutes pulls you right out of it. is there a way to correct or lock in a pronunciation once, so it applies consistently across the rest of that document (or future ones with the same terms)?
VocalVia
@galdayan You’re right — correcting the same term repeatedly would defeat the purpose of an editable workflow.
The proper solution is a reusable pronunciation glossary: add a name, acronym, or technical term once, specify how it should be spoken, preview it, and choose whether the rule applies to one document or the whole workspace. VocalVia would then apply it consistently during synthesis without changing the visible script.
That is not shipped today, so the current workaround is phonetic spelling in the script. But your comment has made the glossary a much clearer priority for us — especially for research papers, where one recurring mispronunciation can ruin the entire listening experience.
VocalVia
@galdayan Hey Gal, quick update: we ended up building this. VocalVia now has a Pronunciation Dictionary in TTS Studio. You can add jargon, acronyms, or author names once and apply the pronunciation to the current document or your whole workspace. It only changes the generated speech, so your original script stays untouched.
Your comment genuinely helped move this to the top of our list. If you try it with a research paper, I’d love to hear where it still falls short.
@zoey_0113 that's a fast turnaround, honestly didn't expect it to jump the queue that quickly. will run a paper with a bunch of author names and acronyms through it and let you know where it holds up. one question - if two documents in the same workspace want the same acronym pronounced differently, does the doc-level rule just override the workspace one, or is there a conflict warning?
VocalVia
@galdayan Great question — the document-level rule takes precedence for that document, while the workspace rule stays unchanged and continues to apply everywhere else.
We don’t show a conflict warning today; the document rule simply acts as an override. A visible “overrides workspace rule” indicator would make that behavior clearer, though, so that’s a useful UI improvement for us.
And yes, please let me know how the author-name and acronym stress test goes!
@zoey_0113 makes sense, override-by-default is the simpler mental model anyway. will try the stress test this week and report back on the acronym pronunciation specifically since that's usually where these tools fall apart.
@Zoey Congrats on the launch of VocalVia! 🚀 Solving the "saved articles I never actually read" problem by turning them into editable multi-voice podcasts is brilliant.
Quick question on handling complex structures: when processing dense documents (like 30+ page research papers with heavy inline citations, tables, or math code), how clean is the initial script translation? Does it automatically distill complex tables into natural conversation, or do you find users usually need to manually edit those segments first?
VocalVia
@franz_briones Thanks, Franz — that’s exactly the kind of challenging document VocalVia is designed to make easier. It first creates an outline and an editable script instead of sending the raw PDF directly to TTS.
For prose and citation-heavy sections, it can usually produce a cleaner spoken structure. I don’t want to overpromise on complex tables or mathematical notation, though — those may still benefit from a quick human review. That’s why every script segment remains editable before audio generation.
I’m continuing to improve this step and would be very interested in testing more real-world research papers.
Editable script before audio is the right call — most TTS tools skip exactly that step. My question is about fiction: in a novel excerpt the speaker is usually implied, not tagged ("she said" disappears after the first exchange). Does the outline step attempt dialogue attribution so characters land on distinct voices, or is fiction a manual-reassignment job today? That feels like the gap between "documents" and "books."
VocalVia
@mystoryland You’ve identified the boundary correctly. Today, fiction still requires manual speaker reassignment because VocalVia works best with explicit roles.
A proper fiction workflow would need a persistent character map, warnings for ambiguous dialogue, and consistent voices across chapters. It’s a direction we’re considering, but I wouldn’t call VocalVia novel-ready yet.
@zoey_0113 Appreciate the honest "not novel-ready yet" — that's rarer than it should be on launch day. Persistent character map + ambiguity warnings is exactly the right frame; cross-chapter consistency is where everything I've tried falls apart. If you ever want a beta tester with a pile of long fiction, happy to break it for you.
VocalVia
@mystoryland Thank you, Olga — that’s an incredibly useful offer. Cross-chapter consistency is exactly the failure mode we would want to test first.
I’ve noted your interest as an early beta tester. When we have a small fiction prototype with a character map and ambiguity review, I’d be glad to reach out before calling it novel-ready.
The multi-voice part caught my eye. When VocalVia turns a document or article into audio, how much control does the user get over which sections use which voice? For example, can someone mark quotes, headings, or different speakers before generation, or is that handled automatically? Since the tagline mentions editable audio, I’m also curious whether edits happen at the text level, the timeline level, or both.
VocalVia
@crystalmei Great question, Xuefei. VocalVia automatically creates an initial multi-speaker script, but users are not locked into those assignments. Before generating the audio, you can edit each segment, change its speaker, and select the voice used by each speaker.
Editing currently happens at the text and segment level rather than on a waveform timeline. This lets you revise the wording and adjust speaker or voice assignments before synthesis. Quotes and sections can be reassigned manually; automatic semantic handling for elements such as headings and quotations is something I’m still improving.
A visual timeline and chapter-level editing would be a valuable next step.
Hey! Editable scripts before audio generation is the part most TTS tools skip, and it's the part that matters. Does it keep speaker assignments stable when you re-edit a segment, or does the outline regenerate?
VocalVia
@vladimir_iudin Yes — the speaker assignment stays attached to the segment when you edit its text. Revising a segment does not regenerate the outline or the rest of the script, so you can iterate on individual lines without losing the existing speaker setup.
The outline or full script is regenerated only when you intentionally return to that earlier step and request a new version.
@zoey_0113 That's the right architecture, segment-level edits without touching the rest. More tools should steal this. Thanks for the detailed answer, good luck with the rest of the launch week.
Multi-voice from a single document is a nice unlock — script studios usually make you assign voices manually shot by shot. How are you handling voice consistency across a long document, especially if the same "character" voice needs to reappear later? That's been one of the trickier problems building voice tooling alongside video generation.
VocalVia
@abhineetarora Voice consistency was one of the reasons I separated speakers from individual script segments. Each speaker keeps the same selected voice throughout the document, and every segment assigned to that speaker uses the same catalog, designed, or cloned voice reference — even when the speaker reappears much later.
The underlying voice identity therefore stays consistent. Prosody and delivery can still vary somewhat between generations, so tighter controls for tone and delivery are an area I’m continuing to improve.
@zoey_0113 That's a solid approach — anchoring consistency at the speaker level rather than per-generation makes sense. The prosody/delivery drift you mentioned is exactly the hard part; we've seen something similar keeping a "character" consistent across multiple video shots, where the underlying identity stays locked but performance nuance shifts each generation. Curious whether you're exploring any kind of reference-conditioning to tighten that, or if it's more of an ongoing model-quality wait-and-see. Congrats on the launch, by the way — following along.