An Empirical Study of Autoregressive Pre-training from Videos
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866929668400087040 |
|---|---|
| author | Rajasegaran, Jathushan Radosavovic, Ilija Ravishankar, Rahul Gandelsman, Yossi Feichtenhofer, Christoph Malik, Jitendra |
| author_facet | Rajasegaran, Jathushan Radosavovic, Ilija Ravishankar, Rahul Gandelsman, Yossi Feichtenhofer, Christoph Malik, Jitendra |
| contents | We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences of visual tokens and train transformer models to autoregressively predict future tokens. Our models are pre-trained on a diverse dataset of videos and images comprising over 1 trillion visual tokens. We explore different architectural, training, and inference design choices. We evaluate the learned visual representations on a range of downstream tasks including image recognition, video classification, object tracking, and robotics. Our results demonstrate that, despite minimal inductive biases, autoregressive pre-training leads to competitive performance across all benchmarks. Finally, we find that scaling our video models results in similar scaling curves to those seen in language models, albeit with a different rate. More details at https://brjathu.github.io/toto/ |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2501_05453 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | An Empirical Study of Autoregressive Pre-training from Videos Rajasegaran, Jathushan Radosavovic, Ilija Ravishankar, Rahul Gandelsman, Yossi Feichtenhofer, Christoph Malik, Jitendra Computer Vision and Pattern Recognition Artificial Intelligence We empirically study autoregressive pre-training from videos. To perform our study, we construct a series of autoregressive video models, called Toto. We treat videos as sequences of visual tokens and train transformer models to autoregressively predict future tokens. Our models are pre-trained on a diverse dataset of videos and images comprising over 1 trillion visual tokens. We explore different architectural, training, and inference design choices. We evaluate the learned visual representations on a range of downstream tasks including image recognition, video classification, object tracking, and robotics. Our results demonstrate that, despite minimal inductive biases, autoregressive pre-training leads to competitive performance across all benchmarks. Finally, we find that scaling our video models results in similar scaling curves to those seen in language models, albeit with a different rate. More details at https://brjathu.github.io/toto/ |
| title | An Empirical Study of Autoregressive Pre-training from Videos |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence |
| url | https://arxiv.org/abs/2501.05453 |