Large Language Models Do Multi-Label Classification Differently

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Ma, Marcus, Chochlakis, Georgios, Pandiyan, Niyantha Maruthu, Thomason, Jesse, Narayanan, Shrikanth
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866912700293971968
author Ma, Marcus
Chochlakis, Georgios
Pandiyan, Niyantha Maruthu
Thomason, Jesse
Narayanan, Shrikanth
author_facet Ma, Marcus
Chochlakis, Georgios
Pandiyan, Niyantha Maruthu
Thomason, Jesse
Narayanan, Shrikanth
contents Multi-label classification is prevalent in real-world settings, but the behavior of Large Language Models (LLMs) in this setting is understudied. We investigate how autoregressive LLMs perform multi-label classification, focusing on subjective tasks, by analyzing the output distributions of the models at each label generation step. We find that the initial probability distribution for the first label often does not reflect the eventual final output, even in terms of relative order and find LLMs tend to suppress all but one label at each generation step. We further observe that as model scale increases, their token distributions exhibit lower entropy and higher single-label confidence, but the internal relative ranking of the labels improves. Finetuning methods such as supervised finetuning and reinforcement learning amplify this phenomenon. We introduce the task of distribution alignment for multi-label settings: aligning LLM-derived label distributions with empirical distributions estimated from annotator responses in subjective tasks. We propose both zero-shot and supervised methods which improve both alignment and predictive performance over existing approaches. We find one method -- taking the max probability over all label generation distributions instead of just using the initial probability distribution -- improves both distribution alignment and overall F1 classification without adding any additional computation.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17510
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Large Language Models Do Multi-Label Classification Differently
Ma, Marcus
Chochlakis, Georgios
Pandiyan, Niyantha Maruthu
Thomason, Jesse
Narayanan, Shrikanth
Computation and Language
Multi-label classification is prevalent in real-world settings, but the behavior of Large Language Models (LLMs) in this setting is understudied. We investigate how autoregressive LLMs perform multi-label classification, focusing on subjective tasks, by analyzing the output distributions of the models at each label generation step. We find that the initial probability distribution for the first label often does not reflect the eventual final output, even in terms of relative order and find LLMs tend to suppress all but one label at each generation step. We further observe that as model scale increases, their token distributions exhibit lower entropy and higher single-label confidence, but the internal relative ranking of the labels improves. Finetuning methods such as supervised finetuning and reinforcement learning amplify this phenomenon. We introduce the task of distribution alignment for multi-label settings: aligning LLM-derived label distributions with empirical distributions estimated from annotator responses in subjective tasks. We propose both zero-shot and supervised methods which improve both alignment and predictive performance over existing approaches. We find one method -- taking the max probability over all label generation distributions instead of just using the initial probability distribution -- improves both distribution alignment and overall F1 classification without adding any additional computation.
title Large Language Models Do Multi-Label Classification Differently
topic Computation and Language
url https://arxiv.org/abs/2505.17510