SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Thoker, Fida Mohammad, Jiang, Letian, Zhao, Chen, Bagad, Piyush, Doughty, Hazel, Ghanem, Bernard, Snoek, Cees G. M.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912316059025408
author Thoker, Fida Mohammad
Jiang, Letian
Zhao, Chen
Bagad, Piyush
Doughty, Hazel
Ghanem, Bernard
Snoek, Cees G. M.
author_facet Thoker, Fida Mohammad
Jiang, Letian
Zhao, Chen
Bagad, Piyush
Doughty, Hazel
Ghanem, Bernard
Snoek, Cees G. M.
contents Continued advances in self-supervised learning have led to significant progress in video representation learning, offering a scalable alternative to supervised approaches by removing the need for manual annotations. Despite strong performance on standard action recognition benchmarks, video self-supervised learning methods are largely evaluated under narrow protocols, typically pretraining on Kinetics-400 and fine-tuning on similar datasets, limiting our understanding of their generalization in real world scenarios. In this work, we present a comprehensive evaluation of modern video self-supervised models, focusing on generalization across four key downstream factors: domain shift, sample efficiency, action granularity, and task diversity. Building on our prior work analyzing benchmark sensitivity in CNN-based contrastive learning, we extend the study to cover state-of-the-art transformer-based video-only and video-text models. Specifically, we benchmark 12 transformer-based methods (7 video-only, 5 video-text) and compare them to 10 CNN-based methods, totaling over 1100 experiments across 8 datasets and 7 downstream tasks. Our analysis shows that, despite architectural advances, transformer-based models remain sensitive to downstream conditions. No method generalizes consistently across all factors, video-only transformers perform better under domain shifts, CNNs outperform for fine-grained tasks, and video-text models often underperform despite large scale pretraining. We also find that recent transformer models do not consistently outperform earlier approaches. Our findings provide a detailed view of the strengths and limitations of current video SSL methods and offer a unified benchmark for evaluating generalization in video representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2504_05706
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning
Thoker, Fida Mohammad
Jiang, Letian
Zhao, Chen
Bagad, Piyush
Doughty, Hazel
Ghanem, Bernard
Snoek, Cees G. M.
Computer Vision and Pattern Recognition
Continued advances in self-supervised learning have led to significant progress in video representation learning, offering a scalable alternative to supervised approaches by removing the need for manual annotations. Despite strong performance on standard action recognition benchmarks, video self-supervised learning methods are largely evaluated under narrow protocols, typically pretraining on Kinetics-400 and fine-tuning on similar datasets, limiting our understanding of their generalization in real world scenarios. In this work, we present a comprehensive evaluation of modern video self-supervised models, focusing on generalization across four key downstream factors: domain shift, sample efficiency, action granularity, and task diversity. Building on our prior work analyzing benchmark sensitivity in CNN-based contrastive learning, we extend the study to cover state-of-the-art transformer-based video-only and video-text models. Specifically, we benchmark 12 transformer-based methods (7 video-only, 5 video-text) and compare them to 10 CNN-based methods, totaling over 1100 experiments across 8 datasets and 7 downstream tasks. Our analysis shows that, despite architectural advances, transformer-based models remain sensitive to downstream conditions. No method generalizes consistently across all factors, video-only transformers perform better under domain shifts, CNNs outperform for fine-grained tasks, and video-text models often underperform despite large scale pretraining. We also find that recent transformer models do not consistently outperform earlier approaches. Our findings provide a detailed view of the strengths and limitations of current video SSL methods and offer a unified benchmark for evaluating generalization in video representation learning.
title SEVERE++: Evaluating Benchmark Sensitivity in Generalization of Video Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.05706