TwelveLabs just introduced Pegasus 1.5, their most significant leap in generative video AI, transforming video into a queryable, structured data asset.
Jockey is the first AI that understands your entire media library just like you do, searching by person, moment, or context across every photo and video you've captured. Powered by TwelveLabs' advanced model stack, Jockey improves automatically with every update. Whether you need to connect via MCP for Claude/ChatGPT or build custom applications using our API, Jockey makes your media library instantly searchable and accessible.
Hey everyone! TwelveLabs just closed a $100M Series B and wanted to share what we're working on for anyone who hasn't come across us before.
We're a video AI company. The core problem we're solving: video is the richest record of reality we have, but machines still can't really understand it. Most systems just convert footage to text and call it a day. We think that's leaving a lot on the table.
Pegasus 1.5 transforms raw video into consistent, structured, timestamped data on-the-fly. Video becomes a queryable and computable asset, based on your company’s custom requirements. Define a schema of what matters in your domain, point it at any video up to 2 hours, and get back structured, time based metadata in a single API call. And, it’s multimodal – pass in an image, and find anytime this reference appears in your video. Your video library, finally queryable for humans and agents.
Marengo 3.0 is TwelveLabs' most significant model to date, delivering human-like video understanding at scale. A multimodal embedding model, Marengo fuses video, audio, and text for holistic video understanding to power precise video search and retrieval.
Rodeo by TwelveLabs is the AI video intelligence platform for creators and teams who produce at scale. Stop wasting hours scrubbing footage. Go from raw clips to a first cut in minutes using plain language. It's structured creation, not manual review. Unlike transcript-first tools, Rodeo's multimodal AI understands visuals, audio, speech, and text simultaneously, making it perfect for visual-first content. Your video library is now instantly queryable for humans and agents.