Long-Tailed Visual Recognition via Permutation-Invariant Head-to-Tail Feature Fusion

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Li, Mengke, Hu, Zhikai, Lu, Yang, Lan, Weichao, Cheung, Yiu-ming, Huang, Hui
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908388015734784
author Li, Mengke
Hu, Zhikai
Lu, Yang
Lan, Weichao
Cheung, Yiu-ming
Huang, Hui
author_facet Li, Mengke
Hu, Zhikai
Lu, Yang
Lan, Weichao
Cheung, Yiu-ming
Huang, Hui
contents The imbalanced distribution of long-tailed data presents a significant challenge for deep learning models, causing them to prioritize head classes while neglecting tail classes. Two key factors contributing to low recognition accuracy are the deformed representation space and a biased classifier, stemming from insufficient semantic information in tail classes. To address these issues, we propose permutation-invariant and head-to-tail feature fusion (PI-H2T), a highly adaptable method. PI-H2T enhances the representation space through permutation-invariant representation fusion (PIF), yielding more clustered features and automatic class margins. Additionally, it adjusts the biased classifier by transferring semantic information from head to tail classes via head-to-tail fusion (H2TF), improving tail class diversity. Theoretical analysis and experiments show that PI-H2T optimizes both the representation space and decision boundaries. Its plug-and-play design ensures seamless integration into existing methods, providing a straightforward path to further performance improvements. Extensive experiments on long-tailed benchmarks confirm the effectiveness of PI-H2T.
format Preprint
id arxiv_https___arxiv_org_abs_2506_00625
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Long-Tailed Visual Recognition via Permutation-Invariant Head-to-Tail Feature Fusion
Li, Mengke
Hu, Zhikai
Lu, Yang
Lan, Weichao
Cheung, Yiu-ming
Huang, Hui
Computer Vision and Pattern Recognition
The imbalanced distribution of long-tailed data presents a significant challenge for deep learning models, causing them to prioritize head classes while neglecting tail classes. Two key factors contributing to low recognition accuracy are the deformed representation space and a biased classifier, stemming from insufficient semantic information in tail classes. To address these issues, we propose permutation-invariant and head-to-tail feature fusion (PI-H2T), a highly adaptable method. PI-H2T enhances the representation space through permutation-invariant representation fusion (PIF), yielding more clustered features and automatic class margins. Additionally, it adjusts the biased classifier by transferring semantic information from head to tail classes via head-to-tail fusion (H2TF), improving tail class diversity. Theoretical analysis and experiments show that PI-H2T optimizes both the representation space and decision boundaries. Its plug-and-play design ensures seamless integration into existing methods, providing a straightforward path to further performance improvements. Extensive experiments on long-tailed benchmarks confirm the effectiveness of PI-H2T.
title Long-Tailed Visual Recognition via Permutation-Invariant Head-to-Tail Feature Fusion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2506.00625