Dual-Encoders for Extreme Multi-Label Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Gupta, Nilesh, Khatri, Devvrit, Rawat, Ankit S, Bhojanapalli, Srinadh, Jain, Prateek, Dhillon, Inderjit
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916162358476800
author Gupta, Nilesh
Khatri, Devvrit
Rawat, Ankit S
Bhojanapalli, Srinadh
Jain, Prateek
Dhillon, Inderjit
author_facet Gupta, Nilesh
Khatri, Devvrit
Rawat, Ankit S
Bhojanapalli, Srinadh
Jain, Prateek
Dhillon, Inderjit
contents Dual-encoder (DE) models are widely used in retrieval tasks, most commonly studied on open QA benchmarks that are often characterized by multi-class and limited training data. In contrast, their performance in multi-label and data-rich retrieval settings like extreme multi-label classification (XMC), remains under-explored. Current empirical evidence indicates that DE models fall significantly short on XMC benchmarks, where SOTA methods linearly scale the number of learnable parameters with the total number of classes (documents in the corpus) by employing per-class classification head. To this end, we first study and highlight that existing multi-label contrastive training losses are not appropriate for training DE models on XMC tasks. We propose decoupled softmax loss - a simple modification to the InfoNCE loss - that overcomes the limitations of existing contrastive losses. We further extend our loss design to a soft top-k operator-based loss which is tailored to optimize top-k prediction performance. When trained with our proposed loss functions, standard DE models alone can match or outperform SOTA methods by up to 2% at Precision@1 even on the largest XMC datasets while being 20x smaller in terms of the number of trainable parameters. This leads to more parameter-efficient and universally applicable solutions for retrieval tasks. Our code and models are publicly available at https://github.com/nilesh2797/dexml.
format Preprint
id arxiv_https___arxiv_org_abs_2310_10636
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Dual-Encoders for Extreme Multi-Label Classification
Gupta, Nilesh
Khatri, Devvrit
Rawat, Ankit S
Bhojanapalli, Srinadh
Jain, Prateek
Dhillon, Inderjit
Machine Learning
Dual-encoder (DE) models are widely used in retrieval tasks, most commonly studied on open QA benchmarks that are often characterized by multi-class and limited training data. In contrast, their performance in multi-label and data-rich retrieval settings like extreme multi-label classification (XMC), remains under-explored. Current empirical evidence indicates that DE models fall significantly short on XMC benchmarks, where SOTA methods linearly scale the number of learnable parameters with the total number of classes (documents in the corpus) by employing per-class classification head. To this end, we first study and highlight that existing multi-label contrastive training losses are not appropriate for training DE models on XMC tasks. We propose decoupled softmax loss - a simple modification to the InfoNCE loss - that overcomes the limitations of existing contrastive losses. We further extend our loss design to a soft top-k operator-based loss which is tailored to optimize top-k prediction performance. When trained with our proposed loss functions, standard DE models alone can match or outperform SOTA methods by up to 2% at Precision@1 even on the largest XMC datasets while being 20x smaller in terms of the number of trainable parameters. This leads to more parameter-efficient and universally applicable solutions for retrieval tasks. Our code and models are publicly available at https://github.com/nilesh2797/dexml.
title Dual-Encoders for Extreme Multi-Label Classification
topic Machine Learning
url https://arxiv.org/abs/2310.10636