Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Siyuan, Tian, Juanxi, Wang, Zedong, Zhang, Luyuan, Liu, Zicheng, Jin, Weiyang, Liu, Yang, Sun, Baigui, Li, Stan Z.
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866914968133173248
author Li, Siyuan
Tian, Juanxi
Wang, Zedong
Zhang, Luyuan
Liu, Zicheng
Jin, Weiyang
Liu, Yang
Sun, Baigui
Li, Stan Z.
author_facet Li, Siyuan
Tian, Juanxi
Wang, Zedong
Zhang, Luyuan
Liu, Zicheng
Jin, Weiyang
Liu, Yang
Sun, Baigui
Li, Stan Z.
contents This paper delves into the interplay between vision backbones and optimizers, unvealing an inter-dependent phenomenon termed \textit{\textbf{b}ackbone-\textbf{o}ptimizer \textbf{c}oupling \textbf{b}ias} (BOCB). We observe that canonical CNNs, such as VGG and ResNet, exhibit a marked co-dependency with SGD families, while recent architectures like ViTs and ConvNeXt share a tight coupling with the adaptive learning rate ones. We further show that BOCB can be introduced by both optimizers and certain backbone designs and may significantly impact the pre-training and downstream fine-tuning of vision models. Through in-depth empirical analysis, we summarize takeaways on recommended optimizers and insights into robust vision backbone architectures. We hope this work can inspire the community to question long-held assumptions on backbones and optimizers, stimulate further explorations, and thereby contribute to more robust vision systems. The source code and models are publicly available at https://bocb-ai.github.io/.
format Preprint
id arxiv_https___arxiv_org_abs_2410_06373
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
Li, Siyuan
Tian, Juanxi
Wang, Zedong
Zhang, Luyuan
Liu, Zicheng
Jin, Weiyang
Liu, Yang
Sun, Baigui
Li, Stan Z.
Computer Vision and Pattern Recognition
Machine Learning
This paper delves into the interplay between vision backbones and optimizers, unvealing an inter-dependent phenomenon termed \textit{\textbf{b}ackbone-\textbf{o}ptimizer \textbf{c}oupling \textbf{b}ias} (BOCB). We observe that canonical CNNs, such as VGG and ResNet, exhibit a marked co-dependency with SGD families, while recent architectures like ViTs and ConvNeXt share a tight coupling with the adaptive learning rate ones. We further show that BOCB can be introduced by both optimizers and certain backbone designs and may significantly impact the pre-training and downstream fine-tuning of vision models. Through in-depth empirical analysis, we summarize takeaways on recommended optimizers and insights into robust vision backbone architectures. We hope this work can inspire the community to question long-held assumptions on backbones and optimizers, stimulate further explorations, and thereby contribute to more robust vision systems. The source code and models are publicly available at https://bocb-ai.github.io/.
title Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
topic Computer Vision and Pattern Recognition
Machine Learning
url https://arxiv.org/abs/2410.06373