Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Siyuan, Tian, Juanxi, Wang, Zedong, Zhang, Luyuan, Liu, Zicheng, Jin, Weiyang, Liu, Yang, Sun, Baigui, Li, Stan Z. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenMixup: Open Mixup Toolbox and Benchmark for Visual Representation Learning
by: Li, Siyuan, et al.
Published: (2022)
by: Li, Siyuan, et al.
Published: (2022)
Switch EMA: A Free Lunch for Better Flatness and Sharpness
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
by: Li, Siyuan, et al.
Published: (2023)
by: Li, Siyuan, et al.
Published: (2023)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
A Survey on Mixup Augmentations and Beyond
by: Jin, Xin, et al.
Published: (2024)
by: Jin, Xin, et al.
Published: (2024)
GenURL: A General Framework for Unsupervised Representation Learning
by: Li, Siyuan, et al.
Published: (2021)
by: Li, Siyuan, et al.
Published: (2021)
Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning
by: Wang, Zedong, et al.
Published: (2025)
by: Wang, Zedong, et al.
Published: (2025)
MogaNet: Multi-order Gated Aggregation Network
by: Li, Siyuan, et al.
Published: (2022)
by: Li, Siyuan, et al.
Published: (2022)
RankE: End-to-End Post-Training for Discrete Text-to-Image Generation with Decoder Co-Evolution
by: Jian, Siyong, et al.
Published: (2026)
by: Jian, Siyong, et al.
Published: (2026)
SemiReward: A General Reward Model for Semi-supervised Learning
by: Li, Siyuan, et al.
Published: (2023)
by: Li, Siyuan, et al.
Published: (2023)
Envision: Benchmarking Unified Understanding & Generation for Causal World Process Insights
by: Tian, Juanxi, et al.
Published: (2025)
by: Tian, Juanxi, et al.
Published: (2025)
MergeTok: Unified Continuous and Discrete Visual Tokenization via Token Merging
by: Zhang, Luyuan, et al.
Published: (2026)
by: Zhang, Luyuan, et al.
Published: (2026)
MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning
by: Li, Zhe, et al.
Published: (2024)
by: Li, Zhe, et al.
Published: (2024)
Taming LLMs by Scaling Learning Rates with Gradient Grouping
by: Li, Siyuan, et al.
Published: (2025)
by: Li, Siyuan, et al.
Published: (2025)
USTEP: Spatio-Temporal Predictive Learning under A Unified View
by: Tan, Cheng, et al.
Published: (2023)
by: Tan, Cheng, et al.
Published: (2023)
PRCL: Probabilistic Representation Contrastive Learning for Semi-Supervised Semantic Segmentation
by: Xie, Haoyu, et al.
Published: (2024)
by: Xie, Haoyu, et al.
Published: (2024)
TopoFR: A Closer Look at Topology Alignment on Face Recognition
by: Dan, Jun, et al.
Published: (2024)
by: Dan, Jun, et al.
Published: (2024)
Unveiling and Mitigating Bias in Audio Visual Segmentation
by: Sun, Peiwen, et al.
Published: (2024)
by: Sun, Peiwen, et al.
Published: (2024)
Revisiting Cross-Attention Mechanisms: Leveraging Beneficial Noise for Domain-Adaptive Learning
by: Zang, Zelin, et al.
Published: (2026)
by: Zang, Zelin, et al.
Published: (2026)
RoboView-Bias: Benchmarking Visual Bias in Embodied Agents for Robotic Manipulation
by: Liu, Enguang, et al.
Published: (2025)
by: Liu, Enguang, et al.
Published: (2025)
Learning Unified Representations from Heterogeneous Data for Robust Heart Rate Modeling
by: Huang, Zhengdong, et al.
Published: (2025)
by: Huang, Zhengdong, et al.
Published: (2025)
Beyond Talking -- Generating Holistic 3D Human Dyadic Motion for Communication
by: Sun, Mingze, et al.
Published: (2024)
by: Sun, Mingze, et al.
Published: (2024)
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
TransFace++: Rethinking the Face Recognition Paradigm with a Focus on Accuracy, Efficiency, and Security
by: Dan, Jun, et al.
Published: (2023)
by: Dan, Jun, et al.
Published: (2023)
LSKNet: A Foundation Lightweight Backbone for Remote Sensing
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
What-Where Transformer: A Slot-Centric Visual Backbone for Concurrent Representation and Localization
by: Yoshihashi, Ryota, et al.
Published: (2026)
by: Yoshihashi, Ryota, et al.
Published: (2026)
Hybrid Global-Local Representation with Augmented Spatial Guidance for Zero-Shot Referring Image Segmentation
by: Liu, Ting, et al.
Published: (2025)
by: Liu, Ting, et al.
Published: (2025)
FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled Audio
by: Xu, Chao, et al.
Published: (2024)
by: Xu, Chao, et al.
Published: (2024)
Towards Efficient and Effective Text-to-Video Retrieval with Coarse-to-Fine Visual Representation Learning
by: Tian, Kaibin, et al.
Published: (2024)
by: Tian, Kaibin, et al.
Published: (2024)
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
by: Li, Puhao, et al.
Published: (2024)
by: Li, Puhao, et al.
Published: (2024)
Short-Long Convolutions Help Hardware-Efficient Linear Attention to Focus on Long Sequences
by: Liu, Zicheng, et al.
Published: (2024)
by: Liu, Zicheng, et al.
Published: (2024)
TP2O: Creative Text Pair-to-Object Generation using Balance Swap-Sampling
by: Li, Jun, et al.
Published: (2023)
by: Li, Jun, et al.
Published: (2023)
OmniFit: Multi-modal 3D Body Fitting via Scale-agnostic Dense Landmark Prediction
by: Cai, Zeyu, et al.
Published: (2026)
by: Cai, Zeyu, et al.
Published: (2026)
BadBone: Backdoor Attacks Against Backbone Models in Visual Prompt Learning
by: Yang, Ziqing, et al.
Published: (2026)
by: Yang, Ziqing, et al.
Published: (2026)
Retrieval Meets Reasoning: Even High-school Textbook Knowledge Benefits Multimodal Reasoning
by: Tan, Cheng, et al.
Published: (2024)
by: Tan, Cheng, et al.
Published: (2024)
Verbalized Machine Learning: Revisiting Machine Learning with Language Models
by: Xiao, Tim Z., et al.
Published: (2024)
by: Xiao, Tim Z., et al.
Published: (2024)
VMDiff: Visual Mixing Diffusion for Limitless Cross-Object Synthesis
by: Xiong, Zeren, et al.
Published: (2025)
by: Xiong, Zeren, et al.
Published: (2025)
Empowering Backbone Models for Visual Text Generation with Input Granularity Control and Glyph-Aware Training
by: Li, Wenbo, et al.
Published: (2024)
by: Li, Wenbo, et al.
Published: (2024)
There is no SAMantics! Exploring SAM as a Backbone for Visual Understanding Tasks
by: Espinosa, Miguel, et al.
Published: (2024)
by: Espinosa, Miguel, et al.
Published: (2024)
Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
by: Su, Zhuo, et al.
Published: (2024)
by: Su, Zhuo, et al.
Published: (2024)
Similar Items
-
OpenMixup: Open Mixup Toolbox and Benchmark for Visual Representation Learning
by: Li, Siyuan, et al.
Published: (2022) -
Switch EMA: A Free Lunch for Better Flatness and Sharpness
by: Li, Siyuan, et al.
Published: (2024) -
Masked Modeling for Self-supervised Representation Learning on Vision and Beyond
by: Li, Siyuan, et al.
Published: (2023) -
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
by: Li, Siyuan, et al.
Published: (2025) -
A Survey on Mixup Augmentations and Beyond
by: Jin, Xin, et al.
Published: (2024)