VocalVia
Turn documents and articles into editable multi-voice audio
111 followers
Turn documents and articles into editable multi-voice audio
111 followers
VocalVia turns PDFs, Word files, Markdown, web articles, and pasted text into structured outlines, editable podcast scripts, and natural multi-voice audio. Choose speakers and voices, refine individual segments, then export the finished audio. It is built for saved reading, study notes, research papers, and long-form content you want to listen to away from a screen.







The editable-per-segment part is what stands out to me over plain TTS — being able to refine one speaker's line without re-rendering the whole thing is the tedious bit everywhere else. Turning a research paper into a two-host format is a genuinely nice use case. How much control do you get over pacing and emphasis within a segment, or is picking the voice the main lever right now?
VocalVia
@lennoxbeflying You’re pointing at exactly the workflow we’re aiming for. Today, you can edit individual lines, assign voices, and add delivery cues such as pauses, emphasis, whispering, sighs, or chuckles. There are also global speed, volume, and loudness controls.
One clarification: generation currently renders the complete audio piece. Regenerating only one changed segment is not shipped yet, but it’s high on the roadmap because it would make iteration much faster.
I believe that editing the script will take time. If I have to review the content anyway, I might just read the PDF instead. Were there users that could relate to my feedback ?
VocalVia
@reda_roqai_chaoui That’s a fair concern. Editing is optional rather than a required review pass: you can accept the generated script and go straight to audio. The editor is there for people who want to change the structure or verify important details.
There’s also an Original mode that keeps the wording and order intact while mainly removing formatting that does not work when spoken. We’re still early, so I don’t want to claim a broad user pattern yet. This is exactly the tradeoff we’re trying to learn more about.
How does it know when to omit or alter details for spoken format compared to keeping them as is, particularly when referring to technical aspects such as reference, numbers, or technical jargon that requires accuracy?
VocalVia
@saksham_salvi The user chooses that tradeoff before generation. Summary may omit minor details, Detailed preserves more context, and Original keeps the wording and order while only removing formatting that does not work when spoken.
For technical documents, Original or Detailed with instructions to preserve references, numbers, units, and terminology is the safest option. We still show the script before synthesis because AI transformation should not be treated as a fidelity guarantee.
A timeline or chapter marker system would be really useful so I can jump back to a specific section in long papers without rewinding through the whole audio.
VocalVia
@tahsinm7m7 That makes a lot of sense, Tahsin. A clickable chapter list with timestamps would make long research audio much easier to navigate, revisit, and resume.
This isn’t available in the current version yet, but I’ve added it to the roadmap. My preferred approach is to generate the initial chapters from the document outline, while still allowing users to rename, reorder, or adjust them before export. Thanks for the very practical suggestion.
The multi-voice output caught me off guard, it actually feels like a real conversation instead of that robotic single narrator you get elsewhere. Refining one segment without regenerating the whole file is a lifesaver for long papers.
VocalVia
@dne4ga7 Thank you, Döne — hearing that is very encouraging.
Long documents make even small mistakes expensive, so keeping the script editable at the segment level is a core part of VocalVia. I’m continuing to improve voice consistency and delivery control while preserving that ability to refine specific parts of the script.
Is it primarily for generating podcasts?
VocalVia
@ragsyme Podcast-style, multi-speaker audio is a core use case, but it isn’t the only one. VocalVia also supports single-voice narration and voiceovers for research papers, saved articles, study notes, tutorials, and other long-form documents.
The common idea is to turn the source material into an editable script first, so users can review the structure, wording, speakers, and voices before generating the final audio.
Finally tried it on a 40-page research paper and the multi-voice narration actually makes the dense sections easier to follow than my usual text-to-speech. Editing individual segments before export is a nice touch too.
VocalVia
@mirag1jy Thank you for putting a full 40-page research paper through it, Mira — that’s exactly the kind of real-world test I was hoping for. I’m especially glad the multi-voice structure made the dense sections easier to follow and that segment editing felt useful.
If you noticed any sections where citations, tables, pronunciation, or pacing still sounded awkward, I’d genuinely appreciate the details. That feedback would be very useful for improving the document-to-script step.