I've been building video tooling for a few months and the thing that still gets me is how confidently a model will describe a video it barely saw.
Most pipelines hand it a frame every few seconds plus a transcript. Then you ask about pacing, or camera work, or whether the speaker sounded unsure, and it answers. Fluently. The answer just isn't grounded in anything that was actually passed in.
Turn any video into agent-ready context: camera moves and pacing, voice emotion, speakers, on-screen text. Inspect it all in a local viewer, or get a written breakdown in one command. Built on the free MIT claude-real-video. $29 once — $19 launch price thru Aug 31.