Gemini gets agentic video understanding
Instead of a fixed frame rate, the model picks which segments to inspect. Google claims up to 88 percent fewer tokens; no outside test confirms it yet.
Symbolic image: at an editing bay, a person seen from behind turns a jog wheel while a monitor wall shows blurred video tiles and a playback machine blinks status lights.
Gemini can now analyze video agentically: the model decides on its own which passages to examine and how closely, instead of sampling every recording at a fixed frame rate.
At a glance
- Available for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite; Google reports the strongest figures for 3.7 Flash.
- Google puts the gains at up to 88 percent fewer tokens, up to 66 percent lower cost and up to 7 percent better accuracy.
- Access runs through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform, covering uploads and YouTube links.
- No extra charge: standard Gemini API token pricing applies.
- The Gemini app is due to get it soon, and YouTube's Ask YouTube in the coming months, Google says.
Google has switched on agentic video understanding in Gemini. Rather than sampling a recording at a fixed frame rate, the model chooses which passages to look at and how closely. The company announced the change on its own blog, starting with three models.
How the processing changes
Conventional video analysis breaks footage into evenly spaced frames. Google describes Gemini as searching the material deliberately instead, moving between the visual frames, the audio and the transcript. How fast it scans follows the question being asked rather than a preset value.
Developers switch the behavior on with a configuration value in the Gemini API; the announcement includes a short Python snippet. Uploaded files and YouTube videos are both supported.
The figures come from Google
Three numbers are cited: up to 88 percent lower token consumption, up to 66 percent lower cost and up to 7 percent better accuracy. Gemini 3.7 Flash is presented as the strongest of the set, with 3.6 Flash and 3.5 Flash-Lite also covered.
The announcement does not identify the test data behind those percentages. They are vendor figures, each carrying an “up to” qualifier, and no independent measurement was available at the time of writing.
What Google says it is for
The post lists sub-second moment retrieval for editing work and needle-in-a-haystack search across multi-hour recordings. It also names anomaly detection, where the model re-scans a suspect stretch more densely, and counting actions or objects over time.
Availability and price
The capability runs through the Gemini API in Google AI Studio and through the Gemini Enterprise Agent Platform. There is no separate feature fee; standard token pricing applies. Google says the Gemini app will receive it soon and that YouTube's Ask YouTube will follow in the coming months, naming no date in either case.
What this report does not establish
Only one source was available for this article: Google's own announcement. The performance claims, the model assignments and the timelines are therefore uncross-checked. Whether short clips see savings comparable to multi-hour footage is not addressed either.
FAQ
What does agentic video understanding mean in Gemini?
The model no longer samples a video at a fixed frame rate; it searches for the relevant segments itself, switching between visuals, audio and the transcript.
Which Gemini models support it?
Per Google: Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite. The company reports its best figures for 3.7 Flash.
Does agentic video understanding cost extra?
No. Standard Gemini API token pricing applies, and Google names no additional fee for the feature.