Video Representation Learning with Joint-Embedding Predictive Architectures

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Drozdov, Katrina, Shwartz-Ziv, Ravid, LeCun, Yann
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909428393967616
author Drozdov, Katrina
Shwartz-Ziv, Ravid
LeCun, Yann
author_facet Drozdov, Katrina
Shwartz-Ziv, Ravid
LeCun, Yann
contents Video representation learning is an increasingly important topic in machine learning research. We present Video JEPA with Variance-Covariance Regularization (VJ-VCR): a joint-embedding predictive architecture for self-supervised video representation learning that employs variance and covariance regularization to avoid representation collapse. We show that hidden representations from our VJ-VCR contain abstract, high-level information about the input data. Specifically, they outperform representations obtained from a generative baseline on downstream tasks that require understanding of the underlying dynamics of moving objects in the videos. Additionally, we explore different ways to incorporate latent variables into the VJ-VCR framework that capture information about uncertainty in the future in non-deterministic settings.
format Preprint
id arxiv_https___arxiv_org_abs_2412_10925
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Video Representation Learning with Joint-Embedding Predictive Architectures
Drozdov, Katrina
Shwartz-Ziv, Ravid
LeCun, Yann
Computer Vision and Pattern Recognition
Artificial Intelligence
Video representation learning is an increasingly important topic in machine learning research. We present Video JEPA with Variance-Covariance Regularization (VJ-VCR): a joint-embedding predictive architecture for self-supervised video representation learning that employs variance and covariance regularization to avoid representation collapse. We show that hidden representations from our VJ-VCR contain abstract, high-level information about the input data. Specifically, they outperform representations obtained from a generative baseline on downstream tasks that require understanding of the underlying dynamics of moving objects in the videos. Additionally, we explore different ways to incorporate latent variables into the VJ-VCR framework that capture information about uncertainty in the future in non-deterministic settings.
title Video Representation Learning with Joint-Embedding Predictive Architectures
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2412.10925