Human AI Agents have a real face and a real voice, and run live conversation rather than turn-based exchange. Interrupt mid-sentence, trail off, talk over it, and the endpointing holds. Setup is one still photo, a persona and a voice. No rig, no capture session, no script. Two face models behind a single API. Portrait for scale, Presence for expressiveness. Both drop into Pipecat and LiveKit. Try to interrupt it. Most demos cannot survive that, and it is the fastest way to judge this one.
I'm very excited to finally have Ojin released!! A lot of engineering efforts went into making this product possible and I'm so proud of the team. Can't wait for people to try it out!
Report
Maker
We're very excited to share Ojin today. Please check it out and share your feedback, thoughts and questions. We're eager to hear what you think - the more feedback, the better!
Report
Turn-taking being most of the product is the right admission, and it's also the thing nobody buys on. We work on the generated side of avatars and the failure mode is different, a face that holds for eight seconds and falls apart on the ninth, but the rule is the same, the boring temporal part is where the work actually is. I'd put endpointing latency on the page next to the face model specs, because two face models behind one API reads like the differentiator and it isn't. Anyone who has shipped voice will pick you on the interrupt handling.
Real-time voice makes $/minute tempting, but I’d rather know cost per completed call outcome. A 4-minute call that books the meeting can be cheaper than a 60-second call that goes nowhere.
Report
@zhangchen That is true and per-minute pricing hides exactly that, a four-minute call that lands beats a fast one that goes nowhere.
I'm very excited to finally have Ojin released!! A lot of engineering efforts went into making this product possible and I'm so proud of the team. Can't wait for people to try it out!
We're very excited to share Ojin today. Please check it out and share your feedback, thoughts and questions. We're eager to hear what you think - the more feedback, the better!
Turn-taking being most of the product is the right admission, and it's also the thing nobody buys on. We work on the generated side of avatars and the failure mode is different, a face that holds for eight seconds and falls apart on the ninth, but the rule is the same, the boring temporal part is where the work actually is. I'd put endpointing latency on the page next to the face model specs, because two face models behind one API reads like the differentiator and it isn't. Anyone who has shipped voice will pick you on the interrupt handling.
DROP
Real-time voice makes $/minute tempting, but I’d rather know cost per completed call outcome. A 4-minute call that books the meeting can be cheaper than a 60-second call that goes nowhere.
@zhangchen That is true and per-minute pricing hides exactly that, a four-minute call that lands beats a fast one that goes nowhere.