Recurrent Video Masked Autoencoders

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Zoran, Daniel, Parthasarathy, Nikhil, Yang, Yi, Hudson, Drew A, Carreira, Joao, Zisserman, Andrew
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913049839927296
author Zoran, Daniel
Parthasarathy, Nikhil
Yang, Yi
Hudson, Drew A
Carreira, Joao
Zisserman, Andrew
author_facet Zoran, Daniel
Parthasarathy, Nikhil
Yang, Yi
Hudson, Drew A
Carreira, Joao
Zisserman, Andrew
contents We present Recurrent Video Masked-Autoencoders (RVM): a novel approach to video representation learning that leverages recurrent computation to model the temporal structure of video data. RVM couples an asymmetric masking objective with a transformer-based recurrent neural network to aggregate information over time, training solely on a simple pixel reconstruction loss. This design yields a highly efficient "generalist" encoder: RVM achieves competitive performance with state-of-the-art video models (e.g. VideoMAE, V-JEPA) on video-level tasks like action classification, and point and object tracking, while matching or exceeding the performance of image models (e.g. DINOv2) on tasks that require strong geometric and dense spatial features. Notably, RVM achieves strong performance in the small-model regime without requiring knowledge distillation, exhibiting up to 30x greater parameter efficiency than competing video masked autoencoders. Finally, we demonstrate that RVM's recurrent nature allows for stable feature propagation over long temporal horizons with linear computational cost, overcoming some of the limitations of standard spatio-temporal attention-based video models. Ablation studies further highlight the factors driving the model's success, with qualitative results showing that RVM learns rich representations of scene semantics, structure, and motion.
format Preprint
id arxiv_https___arxiv_org_abs_2512_13684
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Recurrent Video Masked Autoencoders
Zoran, Daniel
Parthasarathy, Nikhil
Yang, Yi
Hudson, Drew A
Carreira, Joao
Zisserman, Andrew
Computer Vision and Pattern Recognition
We present Recurrent Video Masked-Autoencoders (RVM): a novel approach to video representation learning that leverages recurrent computation to model the temporal structure of video data. RVM couples an asymmetric masking objective with a transformer-based recurrent neural network to aggregate information over time, training solely on a simple pixel reconstruction loss. This design yields a highly efficient "generalist" encoder: RVM achieves competitive performance with state-of-the-art video models (e.g. VideoMAE, V-JEPA) on video-level tasks like action classification, and point and object tracking, while matching or exceeding the performance of image models (e.g. DINOv2) on tasks that require strong geometric and dense spatial features. Notably, RVM achieves strong performance in the small-model regime without requiring knowledge distillation, exhibiting up to 30x greater parameter efficiency than competing video masked autoencoders. Finally, we demonstrate that RVM's recurrent nature allows for stable feature propagation over long temporal horizons with linear computational cost, overcoming some of the limitations of standard spatio-temporal attention-based video models. Ablation studies further highlight the factors driving the model's success, with qualitative results showing that RVM learns rich representations of scene semantics, structure, and motion.
title Recurrent Video Masked Autoencoders
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2512.13684