We just raised $100M here's what we're building
Hey everyone! TwelveLabs just closed a $100M Series B and wanted to share what we're working on for anyone who hasn't come across us before.
We're a video AI company. The core problem we're solving: video is the richest record of reality we have, but machines still can't really understand it. Most systems just convert footage to text and call it a day. We think that's leaving a lot on the table.
Our platform has two main models:
honestly the search is way sharper than i expected, like i threw in a random cooking video and it pulled out the exact moment they mentioned "fold in the eggs" without any tagging. pretty cool to actually see video understanding feel useful
the split between Marengo for retrieval and Pegasus for schema-based segmentation is a real architecture, not just a wrapper on top of an LLM. one practical question - for personal footage across years of different phones/cameras, is there an upfront processing step you have to run before search works, or does it index incrementally as you add new files to the library?
Congrats on the launch. The whole-library search angle is strong, especially for teams with years of product demos, webinars, and raw footage. When Jockey returns a highlight reel or a moment-based answer, how much provenance does the user get back? I would want every suggestion to point to exact source clips and timestamps so the result is easy to verify before editing or publishing.
A live collaboration mode where teams can annotate and tag specific video segments together in real time would be huge. Right now analysis feels like a solo task, but most of our video review happens in group settings where marketers, editors, and strategists need to align on what they see.
honestly the search looks solid, but it would be super helpful if you could save and share specific video clips or search results with timestamps baked in. basically a way to send someone straight to the exact moment in the video rather than just linking the whole thing and saying "go to 4:32". that would make it way more useful for team collaboration
The way the search results show exact timestamps with preview thumbnails makes me feel like I'm scanning a real video library, not just a text index. That attention to temporal precision shows serious craft.
Searched a bunch of old vacation clips by typing "sunset over water" and it actually pulled the right moments, which kind of startled me. The natural language search feels way more useful than tagging everything manually.