Differential-informed Sample Selection Accelerates Multimodal Contrastive Learning

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Zhao, Zihua, Hong, Feng, Chen, Mengxi, Chen, Pengyi, Liu, Benyuan, Yao, Jiangchao, Zhang, Ya, Wang, Yanfeng
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915395335618560
author Zhao, Zihua
Hong, Feng
Chen, Mengxi
Chen, Pengyi
Liu, Benyuan
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
author_facet Zhao, Zihua
Hong, Feng
Chen, Mengxi
Chen, Pengyi
Liu, Benyuan
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
contents The remarkable success of contrastive-learning-based multimodal models has been greatly driven by training on ever-larger datasets with expensive compute consumption. Sample selection as an alternative efficient paradigm plays an important direction to accelerate the training process. However, recent advances on sample selection either mostly rely on an oracle model to offline select a high-quality coreset, which is limited in the cold-start scenarios, or focus on online selection based on real-time model predictions, which has not sufficiently or efficiently considered the noisy correspondence. To address this dilemma, we propose a novel Differential-Informed Sample Selection (DISSect) method, which accurately and efficiently discriminates the noisy correspondence for training acceleration. Specifically, we rethink the impact of noisy correspondence on contrastive learning and propose that the differential between the predicted correlation of the current model and that of a historical model is more informative to characterize sample quality. Based on this, we construct a robust differential-based sample selection and analyze its theoretical insights. Extensive experiments on three benchmark datasets and various downstream tasks demonstrate the consistent superiority of DISSect over current state-of-the-art methods. Source code is available at: https://github.com/MediaBrain-SJTU/DISSect.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12998
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Differential-informed Sample Selection Accelerates Multimodal Contrastive Learning
Zhao, Zihua
Hong, Feng
Chen, Mengxi
Chen, Pengyi
Liu, Benyuan
Yao, Jiangchao
Zhang, Ya
Wang, Yanfeng
Computer Vision and Pattern Recognition
Machine Learning
The remarkable success of contrastive-learning-based multimodal models has been greatly driven by training on ever-larger datasets with expensive compute consumption. Sample selection as an alternative efficient paradigm plays an important direction to accelerate the training process. However, recent advances on sample selection either mostly rely on an oracle model to offline select a high-quality coreset, which is limited in the cold-start scenarios, or focus on online selection based on real-time model predictions, which has not sufficiently or efficiently considered the noisy correspondence. To address this dilemma, we propose a novel Differential-Informed Sample Selection (DISSect) method, which accurately and efficiently discriminates the noisy correspondence for training acceleration. Specifically, we rethink the impact of noisy correspondence on contrastive learning and propose that the differential between the predicted correlation of the current model and that of a historical model is more informative to characterize sample quality. Based on this, we construct a robust differential-based sample selection and analyze its theoretical insights. Extensive experiments on three benchmark datasets and various downstream tasks demonstrate the consistent superiority of DISSect over current state-of-the-art methods. Source code is available at: https://github.com/MediaBrain-SJTU/DISSect.
title Differential-informed Sample Selection Accelerates Multimodal Contrastive Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2507.12998