Dynamic Reflections: Probing Video Representations with Text Alignment

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhu, Tyler, Han, Tengda, Guibas, Leonidas, Pătrăucean, Viorica, Ovsjanikov, Maks
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910005250228224
author Zhu, Tyler
Han, Tengda
Guibas, Leonidas
Pătrăucean, Viorica
Ovsjanikov, Maks
author_facet Zhu, Tyler
Han, Tengda
Guibas, Leonidas
Pătrăucean, Viorica
Ovsjanikov, Maks
contents The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encoders across diverse data types. While significant progress has been made in aligning images with text, the temporal nature of video data remains largely unexplored in this context. In this work, we conduct the first comprehensive study of video-text representation alignment, probing the capabilities of modern video and language encoders. Our findings reveal several key insights. First, we demonstrate that cross-modal alignment highly depends on the richness of both visual (static images vs. multi-frame videos) and text (single caption vs. a collection) data provided at test time, especially when using state-of-the-art video encoders. We propose parametric test-time scaling laws that capture this behavior and show remarkable predictive power against empirical observations. Secondly, we investigate the correlation between semantic alignment and performance on both semantic and non-semantic downstream tasks, providing initial evidence that strong alignment against text encoders may be linked to general-purpose video representation and understanding. Finally, we correlate temporal reasoning with cross-modal alignment providing a challenging test-bed for vision and language models. Overall, our work introduces video-text alignment as an informative zero-shot way to probe the representation power of different encoders for spatio-temporal data. Project page can be found at https://video-prh.github.io/
format Preprint
id arxiv_https___arxiv_org_abs_2511_02767
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Dynamic Reflections: Probing Video Representations with Text Alignment
Zhu, Tyler
Han, Tengda
Guibas, Leonidas
Pătrăucean, Viorica
Ovsjanikov, Maks
Computer Vision and Pattern Recognition
The alignment of representations from different modalities has recently been shown to provide insights on the structural similarities and downstream capabilities of different encoders across diverse data types. While significant progress has been made in aligning images with text, the temporal nature of video data remains largely unexplored in this context. In this work, we conduct the first comprehensive study of video-text representation alignment, probing the capabilities of modern video and language encoders. Our findings reveal several key insights. First, we demonstrate that cross-modal alignment highly depends on the richness of both visual (static images vs. multi-frame videos) and text (single caption vs. a collection) data provided at test time, especially when using state-of-the-art video encoders. We propose parametric test-time scaling laws that capture this behavior and show remarkable predictive power against empirical observations. Secondly, we investigate the correlation between semantic alignment and performance on both semantic and non-semantic downstream tasks, providing initial evidence that strong alignment against text encoders may be linked to general-purpose video representation and understanding. Finally, we correlate temporal reasoning with cross-modal alignment providing a challenging test-bed for vision and language models. Overall, our work introduces video-text alignment as an informative zero-shot way to probe the representation power of different encoders for spatio-temporal data. Project page can be found at https://video-prh.github.io/
title Dynamic Reflections: Probing Video Representations with Text Alignment
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2511.02767