Saved in:
Bibliographic Details
Main Authors: Li, Zhong-Yu, Hu, Yu-Song, Yin, Bo-Wen, Cheng, Ming-Ming
Format: Preprint
Published: 2024
Subjects:
Online Access:https://arxiv.org/abs/2411.15787
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916494044037120
author Li, Zhong-Yu
Hu, Yu-Song
Yin, Bo-Wen
Cheng, Ming-Ming
author_facet Li, Zhong-Yu
Hu, Yu-Song
Yin, Bo-Wen
Cheng, Ming-Ming
contents Vision representation learning, especially self-supervised learning, is pivotal for various vision applications. Ensemble learning has also succeeded in enhancing the performance and robustness of the vision models. However, traditional ensemble strategies are impractical for representation learning, especially self-supervised representation learning that requires large-scale datasets and long schedules. This is because they require k times more training and inference computation costs for an ensemble of k models. Differently, we introduce Multi-Token Enhancing (MTE) that extracts multiple auxiliary tokens simultaneously from a single model to enhance representation learning, while incurring minimal additional training costs and no additional inference costs. These auxiliary tokens, including auxiliary CLS tokens and adaptively pooled tokens, capture complementary information due to their differences. Meanwhile, to address the increase in inference costs, we distill the knowledge acquired by the auxiliary tokens into a global token during pre-training. Consequently, we can discard the auxiliary tokens during inference without incurring additional costs. Our MTE is compatible with various self-supervised loss functions and architectures, consistently improving performances across different downstream tasks. Our source code will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2411_15787
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Multi-Token Enhancing for Vision Representation Learning
Li, Zhong-Yu
Hu, Yu-Song
Yin, Bo-Wen
Cheng, Ming-Ming
Computer Vision and Pattern Recognition
Vision representation learning, especially self-supervised learning, is pivotal for various vision applications. Ensemble learning has also succeeded in enhancing the performance and robustness of the vision models. However, traditional ensemble strategies are impractical for representation learning, especially self-supervised representation learning that requires large-scale datasets and long schedules. This is because they require k times more training and inference computation costs for an ensemble of k models. Differently, we introduce Multi-Token Enhancing (MTE) that extracts multiple auxiliary tokens simultaneously from a single model to enhance representation learning, while incurring minimal additional training costs and no additional inference costs. These auxiliary tokens, including auxiliary CLS tokens and adaptively pooled tokens, capture complementary information due to their differences. Meanwhile, to address the increase in inference costs, we distill the knowledge acquired by the auxiliary tokens into a global token during pre-training. Consequently, we can discard the auxiliary tokens during inference without incurring additional costs. Our MTE is compatible with various self-supervised loss functions and architectures, consistently improving performances across different downstream tasks. Our source code will be made publicly available.
title Multi-Token Enhancing for Vision Representation Learning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2411.15787