Clarity - Mute the room, keep the speaker, in real time

Real-time speech enhancement and target speaker extraction for voice agents. Clarity-1 strips out background noise and other people talking on a live call, so your agent hears only the caller. Streams as audio arrives. Hear the before/after on our site.

Add a comment

Replies

Best
Hey everyone, we're back, this time out of Berlin! Our customers keep asking us for the same thing: make agents work when the caller isn't somewhere quiet. Voice agents work well in quiet rooms. Real callers aren't in quiet rooms. They're in cafés, on busy streets, commuting, often with someone talking right next to them. The agent hears all of it, so the caller gets misheard, or the agent gets interrupted by sounds that were never meant for it. We couldn't find a denoiser we were happy shipping to our customers, so we built one. Clarity-1 does two things on a live audio stream, in real time: Speech enhancement: removes background noise like traffic, trains and café chatter Target speaker extraction: removes other voices too, so your agent only hears the person it's talking to It processes audio as it arrives, so it slots into live calls. On the benchmarks we ran (DNSMOS), Clarity-1 had the highest overall quality score of the models we compared, for both noise removal and target-speaker isolation. The full table and before/after samples are here: Try it on your own audio in the dashboard or via the API: . Your first month is completely free. Two things we'd love from you: What's the noisiest place your users call from? If you find audio where Clarity breaks, send it our way. That's the most useful feedback we can get. This is the first of several models we're releasing over the next few weeks. More soon.

 does this still work reliably if the caller is using a really cheap phone microphone?

   Good question.

hey , do you mind sharing a sample with me?

   The model learned to upsample from 8khz to 24khz. Bad mic quality is part of the model training.

target speaker extraction on a live call is a much harder problem than noise suppression - does Clarity-1 need a short enrollment clip of the caller's voice to lock onto them, or does it figure out "the one talking to the agent" cold on a brand new call with zero prior samples?