Agentic Video Understanding in Gemini - Agentic video analysis for faster, smarter Gemini insights
by•
Agentic video understanding is a new Gemini processing mode (3.7 Flash, 3.6 Flash, 3.5 Flash-Lite) that lets the model decide what to watch, at what speed, and through which modality, instead of a fixed frame rate.
Cuts tokens by up to 88%, cost by up to 66%, boosts accuracy up to 7%, biggest wins on long-form video. Live now via Gemini API in AI Studio and Gemini Enterprise Agent Platform, just set processing to "agentic," standard pricing, no extra fee.

Replies
Agentic video understanding is a new processing mode Google just shipped across Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Shoutout to @rkdoshi and team! :)
Instead of scanning footage at a fixed frame rate, the model actively decides what to watch, at what speed, and through which modality (frames, audio, transcript), fetching only the segments it needs via an internal agentic loop.
What makes it different: it's not a separate product, it's a processing mode toggle inside models you likely already use, with standard token pricing and no added fee.
Key features:
Dynamic scanning: model chooses what to inspect and at what FPS
Up to 88% fewer tokens, up to 66% lower cost, up to 7% better accuracy
Sub-second moment retrieval for precise auto-editing
Needle-in-haystack search across multi-hour video
Anomaly detection via variable FPS resampling
Accurate counting of repeated actions/objects
Who it's for: developers and teams building video search, editing, moderation, or analysis tools on long-form content. Try at Google AI Studio
P.S. I hunt the latest and greatest launches in tech, SaaS and AI, follow to be notified → @rohanrecommends
@rkdoshi @rohanrecommends The token reduction caught my attention most. Processing long videos without scanning every frame could make a big difference.