Review of Extreme Multilabel Classification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Dasgupta, Arpan, Lamba, Preeti, Kushwaha, Ankita, Ravish, Kiran, Katyan, Siddhant, Das, Shrutimoy, Kumar, Pawan
Format: Preprint
Veröffentlicht: 2023
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909611436539904
author Dasgupta, Arpan
Lamba, Preeti
Kushwaha, Ankita
Ravish, Kiran
Katyan, Siddhant
Das, Shrutimoy
Kumar, Pawan
author_facet Dasgupta, Arpan
Lamba, Preeti
Kushwaha, Ankita
Ravish, Kiran
Katyan, Siddhant
Das, Shrutimoy
Kumar, Pawan
contents Extreme multi-label classification or XMLC, is an active area of interest in machine learning. Compared to traditional multi-label classification, here the number of labels is extremely large, hence, the name extreme multi-label classification. Using classical one-versus-all classification does not scale in this case due to large number of labels; the same is true for any other classifier. Embedding labels and features into a lower-dimensional space is a common first step in many XMLC methods. Moreover, other issues include existence of head and tail labels, where tail labels are those that occur in a relatively small number of samples. The existence of tail labels creates issues during embedding. This area has invited application of wide range of approaches ranging from bit compression motivated from compressed sensing, tree based embeddings, deep learning based latent space embedding including using attention weights, linear algebra based embeddings such as SVD, clustering, hashing, to name a few. The community has come up with a useful set of metrics to identify correctly the prediction for head or tail labels.
format Preprint
id arxiv_https___arxiv_org_abs_2302_05971
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Review of Extreme Multilabel Classification
Dasgupta, Arpan
Lamba, Preeti
Kushwaha, Ankita
Ravish, Kiran
Katyan, Siddhant
Das, Shrutimoy
Kumar, Pawan
Machine Learning
Extreme multi-label classification or XMLC, is an active area of interest in machine learning. Compared to traditional multi-label classification, here the number of labels is extremely large, hence, the name extreme multi-label classification. Using classical one-versus-all classification does not scale in this case due to large number of labels; the same is true for any other classifier. Embedding labels and features into a lower-dimensional space is a common first step in many XMLC methods. Moreover, other issues include existence of head and tail labels, where tail labels are those that occur in a relatively small number of samples. The existence of tail labels creates issues during embedding. This area has invited application of wide range of approaches ranging from bit compression motivated from compressed sensing, tree based embeddings, deep learning based latent space embedding including using attention weights, linear algebra based embeddings such as SVD, clustering, hashing, to name a few. The community has come up with a useful set of metrics to identify correctly the prediction for head or tail labels.
title Review of Extreme Multilabel Classification
topic Machine Learning
url https://arxiv.org/abs/2302.05971