Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Orlova, Svetlana, Cavagnero, Niccolò, Dubbelman, Gijs
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914579282395136
author Orlova, Svetlana
Cavagnero, Niccolò
Dubbelman, Gijs
author_facet Orlova, Svetlana
Cavagnero, Niccolò
Dubbelman, Gijs
contents Video foundation models achieve strong performance across many video understanding tasks, but typically require large-scale pre-training on massive video datasets, resulting in substantial data and compute costs. In contrast, modern image foundation models already provide powerful spatial representations. This raises an important question: can competitive video models be built by reusing these spatial representations and pre-training only for temporal reasoning? We take initial steps toward exploring a lightweight training paradigm that freezes a pre-trained image foundation model and trains only a recurrent temporal module to process streaming video. By reusing an image foundation model as a spatial encoder, this approach could significantly reduce the amount of video data and compute required compared to end-to-end video pre-training. In this work, we explore the feasibility of this approach before investing in computing for video pre-training. Our empirical findings across multiple video understanding tasks suggest that strong temporal performance can emerge without large-scale video pre-training, motivating future work on recurrent video foundation models obtained by pre-training a temporal module on top of a frozen image foundation model. Code: https://github.com/tue-mps/towards-video-image-frozen .
format Preprint
id arxiv_https___arxiv_org_abs_2605_19137
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
Orlova, Svetlana
Cavagnero, Niccolò
Dubbelman, Gijs
Computer Vision and Pattern Recognition
Video foundation models achieve strong performance across many video understanding tasks, but typically require large-scale pre-training on massive video datasets, resulting in substantial data and compute costs. In contrast, modern image foundation models already provide powerful spatial representations. This raises an important question: can competitive video models be built by reusing these spatial representations and pre-training only for temporal reasoning? We take initial steps toward exploring a lightweight training paradigm that freezes a pre-trained image foundation model and trains only a recurrent temporal module to process streaming video. By reusing an image foundation model as a spatial encoder, this approach could significantly reduce the amount of video data and compute required compared to end-to-end video pre-training. In this work, we explore the feasibility of this approach before investing in computing for video pre-training. Our empirical findings across multiple video understanding tasks suggest that strong temporal performance can emerge without large-scale video pre-training, motivating future work on recurrent video foundation models obtained by pre-training a temporal module on top of a frozen image foundation model. Code: https://github.com/tue-mps/towards-video-image-frozen .
title Towards Data-Efficient Video Pre-training with Frozen Image Foundation Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2605.19137