Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Yang, Cheng-Fu, Yin, Da, Hu, Wenbo, Ji, Heng, Peng, Nanyun, Zhou, Bolei, Chang, Kai-Wei
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:https://arxiv.org/abs/2411.18651
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909725752295424
author Yang, Cheng-Fu
Yin, Da
Hu, Wenbo
Ji, Heng
Peng, Nanyun
Zhou, Bolei
Chang, Kai-Wei
author_facet Yang, Cheng-Fu
Yin, Da
Hu, Wenbo
Ji, Heng
Peng, Nanyun
Zhou, Bolei
Chang, Kai-Wei
contents Humans recognize objects after observing only a few examples, a remarkable capability enabled by their inherent language understanding of the real-world environment. Developing verbalized and interpretable representation can significantly improve model generalization in low-data settings. In this work, we propose Verbalized Representation Learning (VRL), a novel approach for automatically extracting human-interpretable features for object recognition using few-shot data. Our method uniquely captures inter-class differences and intra-class commonalities in the form of natural language by employing a Vision-Language Model (VLM) to identify key discriminative features between different classes and shared characteristics within the same class. These verbalized features are then mapped to numeric vectors through the VLM. The resulting feature vectors can be further utilized to train and infer with downstream classifiers. Experimental results show that, at the same model scale, VRL achieves a 24% absolute improvement over prior state-of-the-art methods while using 95% less data and a smaller mode. Furthermore, compared to human-labeled attributes, the features learned by VRL exhibit a 20% absolute gain when used for downstream classification tasks. Code is available at: https://github.com/joeyy5588/VRL/tree/main.
format Preprint
id arxiv_https___arxiv_org_abs_2411_18651
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Verbalized Representation Learning for Interpretable Few-Shot Generalization
Yang, Cheng-Fu
Yin, Da
Hu, Wenbo
Ji, Heng
Peng, Nanyun
Zhou, Bolei
Chang, Kai-Wei
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Humans recognize objects after observing only a few examples, a remarkable capability enabled by their inherent language understanding of the real-world environment. Developing verbalized and interpretable representation can significantly improve model generalization in low-data settings. In this work, we propose Verbalized Representation Learning (VRL), a novel approach for automatically extracting human-interpretable features for object recognition using few-shot data. Our method uniquely captures inter-class differences and intra-class commonalities in the form of natural language by employing a Vision-Language Model (VLM) to identify key discriminative features between different classes and shared characteristics within the same class. These verbalized features are then mapped to numeric vectors through the VLM. The resulting feature vectors can be further utilized to train and infer with downstream classifiers. Experimental results show that, at the same model scale, VRL achieves a 24% absolute improvement over prior state-of-the-art methods while using 95% less data and a smaller mode. Furthermore, compared to human-labeled attributes, the features learned by VRL exhibit a 20% absolute gain when used for downstream classification tasks. Code is available at: https://github.com/joeyy5588/VRL/tree/main.
title Verbalized Representation Learning for Interpretable Few-Shot Generalization
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2411.18651