Volver al blog

Síguenos y suscríbete

The Video Data Myth: Why AI Hasn't "Watched" the Internet Yet

AI hasn't watched the internet yet – not due to speed, but token costs and a lack of video labels. Discover the real data bottleneck.

John Agger
John AggerPrincipal Industry Marketing Manager para medios de comunicación y entretenimiento, Fastly

There's a neat explanation going around for why AI still seems to struggle with video compared to text: scraping hasn't made it to video yet, because unlike text, video can't be sped up. A model can ingest a million pages of text in the time it takes a human to read one paragraph, but a 90-minute video, the argument goes, takes a model 90 minutes to "watch," because that's just how long the footage runs.

It's a satisfying theory. And it's backwards.

Models Don't Watch Video, They Sample It

Neither training a model nor analyzing video content is held hostage by the video's runtime. An LLM doesn't have to play the file from start to finish and absorb it in real time. A system such as Gemini samples frames, by default, approximately one per second, processes the audio alongside them, and converts both into tokens the model can work with. Much of that processing can happen in parallel. Sure, longer footage means more work, but a 90-minute video doesn't require 90 minutes to "watch."

The numbers back this up. Google's published overhead for Gemini's video pipeline is roughly a 2-second fixed cost plus ~0.5 seconds of processing per minute of video. Run the math on a 90-minute video and you get well under a minute of actual processing overhead – not 90 minutes (How does Gemini process videos?).

Training is even further removed from "real time": video gets chopped into frames and encoded into features offline, in batches, the same way a text corpus gets tokenized before anyone trains on it. There's no step in the pipeline where a model is sitting there, watching.

So if Intake Speed isn't the Bottleneck, What Is?

Perhaps unsurprisingly, the real cost isn't time, but tokens. In other words: the problem is how much it costs to represent, not how long it takes to ingest.

Per Google's own pricing breakdown, an hour of video at low resolution (measured in tokens, not pixel count) runs to roughly 475,000 tokens: about 360,000 for frames and 115,000 for audio (Understand and count tokens). Compare that to an hour's worth of spoken-word transcript, which might land somewhere in the 10,000–15,000 token range.

The gap exists because video frames are enormously redundant: the 40th frame of someone talking usually looks almost identical to the 39th, so a huge share of those tokens are encoding near-duplicate information. Google's recent move to "agentic" video processing, which lets a model select only the relevant portions of a video rather than every frame, cut token usage by up to 88% precisely by targeting this redundancy (Google cuts Gemini video analysis tokens by up to 88%). That's not a speed fix - it's a cost fix.

The Other Real Bottleneck: Nobody Labeled It

Text has a quirk that makes it uniquely easy to scrape for training: it's already self-describing. A paragraph of text is its own label. Video has no equivalent. While some metadata may exist, a clip of someone picking up a cup and setting it down doesn't come with that explanation attached. The vast majority of online video is captioned, if at all, only at the level of "a man in a kitchen," with no annotation of the action, causality, or physics actually happening in the frame.

That's a labeling problem, not a speed problem, and it's the reason labs have shifted toward self-supervised approaches that don't need hand-labeled video at all. It's the frontier right now – not faster ingestion, but new architectures that sidestep the need for labels in the first place.

Why this is Happening Now

If video weren't throttled by playback speed, why is the industry suddenly racing toward it? Because text is running out. One widely cited estimate puts the growth of high-quality, scrapable web text at under 10% a year, while training-set appetite roughly doubles annually – with the two lines projected to cross around 2028 (The data wall is important). Video, unlabeled, redundant, expensive per token, but available in effectively unlimited supply, is the obvious next frontier once text supply tightens. Yann LeCun has made a version of this argument for years: a four-year-old has absorbed more raw sensory data than any LLM's text training set, almost entirely through vision (The Most Expensive Disagreement in AI).

The Internet’s Most Expensive Binge Watch

AI hasn't fully "gotten to" video – but not because intake is slow. It's because video is costly per unit of useful signal (huge, redundant token counts relative to the information it carries) and poorly labeled for the temporal and causal structure that actually matters. The pivot toward video happening right now isn't a sign that someone finally solved a speed problem. It's a sign that the easy source, text, is starting to run dry, and the next hard problem is squeezing signal out of a format that.

Discover how Fastly's edge cloud platform can help you optimize performance and scale data-heavy workloads.

¿Listo para empezar?

Ponte en contacto con nosotros