Transcribing overlapping speakers is a long-standing challenge for AI models.
To date, the approach has been to first transcribe as though there is just one speaker, and then identify and separate the speaker transcripts using speaker identification and diarisation models. This is tricky for overlapping speech.
Chorus v1 is a single model capable of transcribing one speaker at a time. It works by training a separate token for the first speaker to speak and the second speaker to speak. I'm not sure why this wasn't tried before, or at least isn't used more widely (lmk if you find a reference doing it).
Chorus v1 is now live as an open-weights model on HuggingFace and available for use via Trelis Router.
Report
No reviews yetBe the first to leave a review for Chorus v1
Research Buddy