We Need to Talk About Classification Evaluation Metrics in NLP

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Vickers, Peter, Barrault, Loïc, Monti, Emilio, Aletras, Nikolaos
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866916084088569856
author Vickers, Peter
Barrault, Loïc
Monti, Emilio
Aletras, Nikolaos
author_facet Vickers, Peter
Barrault, Loïc
Monti, Emilio
Aletras, Nikolaos
contents In Natural Language Processing (NLP) classification tasks such as topic categorisation and sentiment analysis, model generalizability is generally measured with standard metrics such as Accuracy, F-Measure, or AUC-ROC. The diversity of metrics, and the arbitrariness of their application suggest that there is no agreement within NLP on a single best metric to use. This lack suggests there has not been sufficient examination of the underlying heuristics which each metric encodes. To address this we compare several standard classification metrics with more 'exotic' metrics and demonstrate that a random-guess normalised Informedness metric is a parsimonious baseline for task performance. To show how important the choice of metric is, we perform extensive experiments on a wide range of NLP tasks including a synthetic scenario, natural language understanding, question answering and machine translation. Across these tasks we use a superset of metrics to rank models and find that Informedness best captures the ideal model characteristics. Finally, we release a Python implementation of Informedness following the SciKitLearn classifier format.
format Preprint
id arxiv_https___arxiv_org_abs_2401_03831
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle We Need to Talk About Classification Evaluation Metrics in NLP
Vickers, Peter
Barrault, Loïc
Monti, Emilio
Aletras, Nikolaos
Computation and Language
Machine Learning
In Natural Language Processing (NLP) classification tasks such as topic categorisation and sentiment analysis, model generalizability is generally measured with standard metrics such as Accuracy, F-Measure, or AUC-ROC. The diversity of metrics, and the arbitrariness of their application suggest that there is no agreement within NLP on a single best metric to use. This lack suggests there has not been sufficient examination of the underlying heuristics which each metric encodes. To address this we compare several standard classification metrics with more 'exotic' metrics and demonstrate that a random-guess normalised Informedness metric is a parsimonious baseline for task performance. To show how important the choice of metric is, we perform extensive experiments on a wide range of NLP tasks including a synthetic scenario, natural language understanding, question answering and machine translation. Across these tasks we use a superset of metrics to rank models and find that Informedness best captures the ideal model characteristics. Finally, we release a Python implementation of Informedness following the SciKitLearn classifier format.
title We Need to Talk About Classification Evaluation Metrics in NLP
topic Computation and Language
Machine Learning
url https://arxiv.org/abs/2401.03831