Revisiting Hierarchical Text Classification: Inference and Metrics

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Plaud, Roman, Labeau, Matthieu, Saillenfest, Antoine, Bonald, Thomas
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912068914905088
author Plaud, Roman
Labeau, Matthieu
Saillenfest, Antoine
Bonald, Thomas
author_facet Plaud, Roman
Labeau, Matthieu
Saillenfest, Antoine
Bonald, Thomas
contents Hierarchical text classification (HTC) is the task of assigning labels to a text within a structured space organized as a hierarchy. Recent works treat HTC as a conventional multilabel classification problem, therefore evaluating it as such. We instead propose to evaluate models based on specifically designed hierarchical metrics and we demonstrate the intricacy of metric choice and prediction inference method. We introduce a new challenging dataset and we evaluate fairly, recent sophisticated models, comparing them with a range of simple but strong baselines, including a new theoretically motivated loss. Finally, we show that those baselines are very often competitive with the latest models. This highlights the importance of carefully considering the evaluation methodology when proposing new methods for HTC. Code implementation and dataset are available at \url{https://github.com/RomanPlaud/revisitingHTC}.
format Preprint
id arxiv_https___arxiv_org_abs_2410_01305
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Revisiting Hierarchical Text Classification: Inference and Metrics
Plaud, Roman
Labeau, Matthieu
Saillenfest, Antoine
Bonald, Thomas
Computation and Language
Machine Learning
Hierarchical text classification (HTC) is the task of assigning labels to a text within a structured space organized as a hierarchy. Recent works treat HTC as a conventional multilabel classification problem, therefore evaluating it as such. We instead propose to evaluate models based on specifically designed hierarchical metrics and we demonstrate the intricacy of metric choice and prediction inference method. We introduce a new challenging dataset and we evaluate fairly, recent sophisticated models, comparing them with a range of simple but strong baselines, including a new theoretically motivated loss. Finally, we show that those baselines are very often competitive with the latest models. This highlights the importance of carefully considering the evaluation methodology when proposing new methods for HTC. Code implementation and dataset are available at \url{https://github.com/RomanPlaud/revisitingHTC}.
title Revisiting Hierarchical Text Classification: Inference and Metrics
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2410.01305