Enhancing Features in Long-tailed Data Using Large Vision Model

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Han, Pengxiao, Ye, Changkun, Tong, Jinguang, Jiang, Cuicui, Hong, Jie, Fang, Li, Li, Xuesong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911426288812032
author Han, Pengxiao
Ye, Changkun
Tong, Jinguang
Jiang, Cuicui
Hong, Jie
Fang, Li
Li, Xuesong
author_facet Han, Pengxiao
Ye, Changkun
Tong, Jinguang
Jiang, Cuicui
Hong, Jie
Fang, Li
Li, Xuesong
contents Language-based foundation models, such as large language models (LLMs) or large vision-language models (LVLMs), have been widely studied in long-tailed recognition. However, the need for linguistic data is not applicable to all practical tasks. In this study, we aim to explore using large vision models (LVMs) or visual foundation models (VFMs) to enhance long-tailed data features without any language information. Specifically, we extract features from the LVM and fuse them with features in the baseline network's map and latent space to obtain the augmented features. Moreover, we design several prototype-based losses in the latent space to further exploit the potential of the augmented features. In the experimental section, we validate our approach on two benchmark datasets: ImageNet-LT and iNaturalist2018.
format Preprint
id arxiv_https___arxiv_org_abs_2504_10852
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Enhancing Features in Long-tailed Data Using Large Vision Model
Han, Pengxiao
Ye, Changkun
Tong, Jinguang
Jiang, Cuicui
Hong, Jie
Fang, Li
Li, Xuesong
Computer Vision and Pattern Recognition
Language-based foundation models, such as large language models (LLMs) or large vision-language models (LVLMs), have been widely studied in long-tailed recognition. However, the need for linguistic data is not applicable to all practical tasks. In this study, we aim to explore using large vision models (LVMs) or visual foundation models (VFMs) to enhance long-tailed data features without any language information. Specifically, we extract features from the LVM and fuse them with features in the baseline network's map and latent space to obtain the augmented features. Moreover, we design several prototype-based losses in the latent space to further exploit the potential of the augmented features. In the experimental section, we validate our approach on two benchmark datasets: ImageNet-LT and iNaturalist2018.
title Enhancing Features in Long-tailed Data Using Large Vision Model
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2504.10852