Jockey by TwelveLabs - The video AI agent that understands your whole library
by•
Jockey is the first AI that understands your entire media library just like you do, searching by person, moment, or context across every photo and video you've captured. Powered by TwelveLabs' advanced model stack, Jockey improves automatically with every update. Whether you need to connect via MCP for Claude/ChatGPT or build custom applications using our API, Jockey makes your media library instantly searchable and accessible.


Replies
The way the search results show exact timestamps with preview thumbnails makes me feel like I'm scanning a real video library, not just a text index. That attention to temporal precision shows serious craft.
Searched a bunch of old vacation clips by typing "sunset over water" and it actually pulled the right moments, which kind of startled me. The natural language search feels way more useful than tagging everything manually.
finally a video ai that actually finds the moment i describe instead of just dumping timestamps. tried it on some old footage and the semantic search picked out exactly the scene i was thinking of.
A browser extension that lets you right-click any video and instantly get a chapter breakdown or summary using your Pegasus model would be huge, especially for longer YouTube content or lectures where I don't always want to scrub manually.
honestly the search is way better than i expected, threw in some random clips and it actually pulled out the exact scenes i was thinking of. pretty wild that you can basically ask it questions about what's happening in the video.
honestly the search across hours of footage feels almost scary good, like you throw in a rough idea and it actually pulls the right moments out without much fuss.
One thing that would make this way more useful for me: a timeline-based search view where I can scrub to the exact moment a concept appears, instead of just getting text hits. Right now it sounds like the results are summaries, but for video editing workflows I really need frame-accurate jumps. Would love to see that built in.
Honestly impressed by how well it picks up on subtle visual cues in long videos, not just obvious keywords. Searched a 40 minute documentary for a specific gesture and it nailed it in seconds.
A live collaboration mode would be huge, letting a team tag and comment on different timestamps together while the AI pulls those notes into a shared summary report.
Finally gave this a spin on a few clips and the semantic search actually nails what's happening in the scene, not just matching keywords. Impressed it picked up on subtle actions without me tagging anything.