Revisiting Feature Prediction for Learning Visual Representations from Video
Fuente:
arXiv
Salvato in:
| Autori principali: | , , , , , , , |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
| _version_ | 1866913312022724608 |
|---|---|
| author | Bardes, Adrien Garrido, Quentin Ponce, Jean Chen, Xinlei Rabbat, Michael LeCun, Yann Assran, Mahmoud Ballas, Nicolas |
| author_facet | Bardes, Adrien Garrido, Quentin Ponce, Jean Chen, Xinlei Rabbat, Michael LeCun, Yann Assran, Mahmoud Ballas, Nicolas |
| contents | This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision. The models are trained on 2 million videos collected from public datasets and are evaluated on downstream image and video tasks. Our results show that learning by predicting video features leads to versatile visual representations that perform well on both motion and appearance-based tasks, without adaption of the model's parameters; e.g., using a frozen backbone. Our largest model, a ViT-H/16 trained only on videos, obtains 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2404_08471 |
| institution | arXiv |
| publishDate | 2024 |
| record_format | arxiv |
| spellingShingle | Revisiting Feature Prediction for Learning Visual Representations from Video Bardes, Adrien Garrido, Quentin Ponce, Jean Chen, Xinlei Rabbat, Michael LeCun, Yann Assran, Mahmoud Ballas, Nicolas Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision. The models are trained on 2 million videos collected from public datasets and are evaluated on downstream image and video tasks. Our results show that learning by predicting video features leads to versatile visual representations that perform well on both motion and appearance-based tasks, without adaption of the model's parameters; e.g., using a frozen backbone. Our largest model, a ViT-H/16 trained only on videos, obtains 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K. |
| title | Revisiting Feature Prediction for Learning Visual Representations from Video |
| topic | Computer Vision and Pattern Recognition Artificial Intelligence Machine Learning |
| url | https://arxiv.org/abs/2404.08471 |