We just raised $100M here's what we're building
Hey everyone! TwelveLabs just closed a $100M Series B and wanted to share what we're working on for anyone who hasn't come across us before.
We're a video AI company. The core problem we're solving: video is the richest record of reality we have, but machines still can't really understand it. Most systems just convert footage to text and call it a day. We think that's leaving a lot on the table.
Our platform has two main models:
Turning raw teleoperation and wearable footage into labeled, timestamped robot training data solves a bottleneck that physical AI teams feel constantly, since collecting the video is easy and annotating it is not. Understanding first person perspective specifically is the right specialization. Do the persistent metadata labels stay consistent across long recording sessions?
Spot on. IMHO this could be a game changer for robotics and physical AI teams.
@karimbenkeroum
Hi Kareem, great question! Pegasus supports Time Based Metadata mode which allows for fine atomic action segmentation on long videos. Each segment is then labeled for fields such as action description, hands used, action per hand, tool used, and so on. So you can see a high quality dense captioning output that is perfect for downstream physical AI tasks!
oh yes, and try it for yourself on the playground at twelvelabs.io