An Empirical Study of Autoregressive Pre-training from Videos

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Rajasegaran, Jathushan, Radosavovic, Ilija, Ravishankar, Rahul, Gandelsman, Yossi, Feichtenhofer, Christoph, Malik, Jitendra
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866929668400087040
author Rajasegaran, Jathushan
Radosavovic, Ilija
Ravishankar, Rahul
Gandelsman, Yossi
Feichtenhofer, Christoph
Malik, Jitendra
author_facet Rajasegaran, Jathushan
Radosavovic, Ilija
Ravishankar, Rahul
Gandelsman, Yossi
Feichtenhofer, Christoph
Malik, Jitendra
contents We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences of visual tokens and train transformer models to autoregressively predict future tokens. Our models are pre-trained on a diverse dataset of videos and images comprising over 1 trillion visual tokens. We explore different architectural, training, and inference design choices. We evaluate the learned visual representations on a range of downstream tasks including image recognition, video classification, object tracking, and robotics. Our results demonstrate that, despite minimal inductive biases, autoregressive pre-training leads to competitive performance across all benchmarks. Finally, we find that scaling our video models results in similar scaling curves to those seen in language models, albeit with a different rate. More details at https://brjathu.github.io/toto/
format Preprint
id arxiv_https___arxiv_org_abs_2501_05453
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle An Empirical Study of Autoregressive Pre-training from Videos
Rajasegaran, Jathushan
Radosavovic, Ilija
Ravishankar, Rahul
Gandelsman, Yossi
Feichtenhofer, Christoph
Malik, Jitendra
Computer Vision and Pattern Recognition
Artificial Intelligence
We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences of visual tokens and train transformer models to autoregressively predict future tokens. Our models are pre-trained on a diverse dataset of videos and images comprising over 1 trillion visual tokens. We explore different architectural, training, and inference design choices. We evaluate the learned visual representations on a range of downstream tasks including image recognition, video classification, object tracking, and robotics. Our results demonstrate that, despite minimal inductive biases, autoregressive pre-training leads to competitive performance across all benchmarks. Finally, we find that scaling our video models results in similar scaling curves to those seen in language models, albeit with a different rate. More details at https://brjathu.github.io/toto/
title An Empirical Study of Autoregressive Pre-training from Videos
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2501.05453