Saved in:
Bibliographic Details
Main Authors: Zhuang, Weijun, Huang, Yuqing, Meng, Weikang, Li, Xin, Liu, Ming, Hong, Xiaopeng, Wang, Yaowei, Zuo, Wangmeng
Format: Preprint
Published: 2026
Subjects:
Online Access:https://arxiv.org/abs/2603.22953
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912980223918080
author Zhuang, Weijun
Huang, Yuqing
Meng, Weikang
Li, Xin
Liu, Ming
Hong, Xiaopeng
Wang, Yaowei
Zuo, Wangmeng
author_facet Zhuang, Weijun
Huang, Yuqing
Meng, Weikang
Li, Xin
Liu, Ming
Hong, Xiaopeng
Wang, Yaowei
Zuo, Wangmeng
contents Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer from two fundamental limitations: severe visual information loss under high masking ratios and temporal information leakage caused by inter-frame correlations. To address these challenges, we propose ClusterSTM, a Cluster-Wise Spatio-Temporal Masking strategy for efficient video-language pretraining. ClusterSTM first performs intra-frame clustering to partition visual tokens into multiple semantically independent clusters, then conducts cluster-wise masking by retaining the token with the highest temporal density within each cluster. Our masking strategy ensure that the retained tokens capture holistic video content while exhibit strong temporal correlation. Additionally, we introduce a video-text relevance reconstruction objective that aligns high-level multimodal semantics beyond conventional visual reconstruction. Extensive experiments across multiple benchmarks demonstrate that ClusterSTM achieves superior performance on video-text retrieval, video question answering, and video captioning tasks, establishing a new state-of-the-art among efficient video-language models.
format Preprint
id arxiv_https___arxiv_org_abs_2603_22953
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
Zhuang, Weijun
Huang, Yuqing
Meng, Weikang
Li, Xin
Liu, Ming
Hong, Xiaopeng
Wang, Yaowei
Zuo, Wangmeng
Computer Vision and Pattern Recognition
Large-scale video-language pretraining enables strong generalization across multimodal tasks but often incurs prohibitive computational costs. Although recent advances in masked visual modeling help mitigate this issue, they still suffer from two fundamental limitations: severe visual information loss under high masking ratios and temporal information leakage caused by inter-frame correlations. To address these challenges, we propose ClusterSTM, a Cluster-Wise Spatio-Temporal Masking strategy for efficient video-language pretraining. ClusterSTM first performs intra-frame clustering to partition visual tokens into multiple semantically independent clusters, then conducts cluster-wise masking by retaining the token with the highest temporal density within each cluster. Our masking strategy ensure that the retained tokens capture holistic video content while exhibit strong temporal correlation. Additionally, we introduce a video-text relevance reconstruction objective that aligns high-level multimodal semantics beyond conventional visual reconstruction. Extensive experiments across multiple benchmarks demonstrate that ClusterSTM achieves superior performance on video-text retrieval, video question answering, and video captioning tasks, establishing a new state-of-the-art among efficient video-language models.
title Cluster-Wise Spatio-Temporal Masking for Efficient Video-Language Pretraining
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2603.22953