Generalized Category Discovery with Large Language Models in the Loop

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: An, Wenbin, Shi, Wenkai, Tian, Feng, Lin, Haonan, Wang, QianYing, Wu, Yaqiang, Cai, Mingxiang, Wang, Luyan, Chen, Yan, Zhu, Haiping, Chen, Ping
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913363376734208
author An, Wenbin
Shi, Wenkai
Tian, Feng
Lin, Haonan
Wang, QianYing
Wu, Yaqiang
Cai, Mingxiang
Wang, Luyan
Chen, Yan
Zhu, Haiping
Chen, Ping
author_facet An, Wenbin
Shi, Wenkai
Tian, Feng
Lin, Haonan
Wang, QianYing
Wu, Yaqiang
Cai, Mingxiang
Wang, Luyan
Chen, Yan
Zhu, Haiping
Chen, Ping
contents Generalized Category Discovery (GCD) is a crucial task that aims to recognize both known and novel categories from a set of unlabeled data by utilizing a few labeled data with only known categories. Due to the lack of supervision and category information, current methods usually perform poorly on novel categories and struggle to reveal semantic meanings of the discovered clusters, which limits their applications in the real world. To mitigate the above issues, we propose Loop, an end-to-end active-learning framework that introduces Large Language Models (LLMs) into the training loop, which can boost model performance and generate category names without relying on any human efforts. Specifically, we first propose Local Inconsistent Sampling (LIS) to select samples that have a higher probability of falling to wrong clusters, based on neighborhood prediction consistency and entropy of cluster assignment probabilities. Then we propose a Scalable Query strategy to allow LLMs to choose true neighbors of the selected samples from multiple candidate samples. Based on the feedback from LLMs, we perform Refined Neighborhood Contrastive Learning (RNCL) to pull samples and their neighbors closer to learn clustering-friendly representations. Finally, we select representative samples from clusters corresponding to novel categories to allow LLMs to generate category names for them. Extensive experiments on three benchmark datasets show that Loop outperforms SOTA models by a large margin and generates accurate category names for the discovered clusters. Code and data are available at https://github.com/Lackel/LOOP.
format Preprint
id arxiv_https___arxiv_org_abs_2312_10897
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Generalized Category Discovery with Large Language Models in the Loop
An, Wenbin
Shi, Wenkai
Tian, Feng
Lin, Haonan
Wang, QianYing
Wu, Yaqiang
Cai, Mingxiang
Wang, Luyan
Chen, Yan
Zhu, Haiping
Chen, Ping
Computation and Language
Artificial Intelligence
Machine Learning
Generalized Category Discovery (GCD) is a crucial task that aims to recognize both known and novel categories from a set of unlabeled data by utilizing a few labeled data with only known categories. Due to the lack of supervision and category information, current methods usually perform poorly on novel categories and struggle to reveal semantic meanings of the discovered clusters, which limits their applications in the real world. To mitigate the above issues, we propose Loop, an end-to-end active-learning framework that introduces Large Language Models (LLMs) into the training loop, which can boost model performance and generate category names without relying on any human efforts. Specifically, we first propose Local Inconsistent Sampling (LIS) to select samples that have a higher probability of falling to wrong clusters, based on neighborhood prediction consistency and entropy of cluster assignment probabilities. Then we propose a Scalable Query strategy to allow LLMs to choose true neighbors of the selected samples from multiple candidate samples. Based on the feedback from LLMs, we perform Refined Neighborhood Contrastive Learning (RNCL) to pull samples and their neighbors closer to learn clustering-friendly representations. Finally, we select representative samples from clusters corresponding to novel categories to allow LLMs to generate category names for them. Extensive experiments on three benchmark datasets show that Loop outperforms SOTA models by a large margin and generates accurate category names for the discovered clusters. Code and data are available at https://github.com/Lackel/LOOP.
title Generalized Category Discovery with Large Language Models in the Loop
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2312.10897