Revisiting Feature Prediction for Learning Visual Representations from Video

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bardes, Adrien, Garrido, Quentin, Ponce, Jean, Chen, Xinlei, Rabbat, Michael, LeCun, Yann, Assran, Mahmoud, Ballas, Nicolas
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913312022724608
author Bardes, Adrien
Garrido, Quentin
Ponce, Jean
Chen, Xinlei
Rabbat, Michael
LeCun, Yann
Assran, Mahmoud
Ballas, Nicolas
author_facet Bardes, Adrien
Garrido, Quentin
Ponce, Jean
Chen, Xinlei
Rabbat, Michael
LeCun, Yann
Assran, Mahmoud
Ballas, Nicolas
contents This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision. The models are trained on 2 million videos collected from public datasets and are evaluated on downstream image and video tasks. Our results show that learning by predicting video features leads to versatile visual representations that perform well on both motion and appearance-based tasks, without adaption of the model's parameters; e.g., using a frozen backbone. Our largest model, a ViT-H/16 trained only on videos, obtains 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K.
format Preprint
id arxiv_https___arxiv_org_abs_2404_08471
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting Feature Prediction for Learning Visual Representations from Video
Bardes, Adrien
Garrido, Quentin
Ponce, Jean
Chen, Xinlei
Rabbat, Michael
LeCun, Yann
Assran, Mahmoud
Ballas, Nicolas
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
This paper explores feature prediction as a stand-alone objective for unsupervised learning from video and introduces V-JEPA, a collection of vision models trained solely using a feature prediction objective, without the use of pretrained image encoders, text, negative examples, reconstruction, or other sources of supervision. The models are trained on 2 million videos collected from public datasets and are evaluated on downstream image and video tasks. Our results show that learning by predicting video features leads to versatile visual representations that perform well on both motion and appearance-based tasks, without adaption of the model's parameters; e.g., using a frozen backbone. Our largest model, a ViT-H/16 trained only on videos, obtains 81.9% on Kinetics-400, 72.2% on Something-Something-v2, and 77.9% on ImageNet1K.
title Revisiting Feature Prediction for Learning Visual Representations from Video
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2404.08471