Agentic Video Understanding in Gemini - Agentic video analysis for faster, smarter Gemini insights

by
Agentic video understanding is a new Gemini processing mode (3.7 Flash, 3.6 Flash, 3.5 Flash-Lite) that lets the model decide what to watch, at what speed, and through which modality, instead of a fixed frame rate. Cuts tokens by up to 88%, cost by up to 66%, boosts accuracy up to 7%, biggest wins on long-form video. Live now via Gemini API in AI Studio and Gemini Enterprise Agent Platform, just set processing to "agentic," standard pricing, no extra fee.

Add a comment

Replies

Best
Hunter
📌

Agentic video understanding is a new processing mode Google just shipped across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Shoutout to and team! :)

Instead of scanning footage at a fixed frame rate, the model actively decides what to watch, at what speed, and through which modality (frames, audio, transcript), fetching only the segments it needs via an internal agentic loop.

What makes it different: it's not a separate product, it's a processing mode toggle inside models you likely already use, with standard token pricing and no added fee.

Key features:

  • Dynamic scanning: model chooses what to inspect and at what FPS

  • Up to 88% fewer tokens, up to 66% lower cost, up to 7% better accuracy

  • Sub-second moment retrieval for precise auto-editing

  • Needle-in-haystack search across multi-hour video

  • Anomaly detection via variable FPS resampling

  • Accurate counting of repeated actions/objects

Who it's for: developers and teams building video search, editing, moderation, or analysis tools on long-form content. Try at

P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified

   The token reduction caught my attention most. Processing long videos without scanning every frame could make a big difference.

the percentage cuts in tokens and costs caught my attention for real. And at that speed?