To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gonsior, Julius, Falkenberg, Christian, Magino, Silvio, Reusch, Anja, Thiele, Maik, Lehner, Wolfgang
Format: Preprint
Veröffentlicht: 2022
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866912270244642816
author Gonsior, Julius
Falkenberg, Christian
Magino, Silvio
Reusch, Anja
Thiele, Maik
Lehner, Wolfgang
author_facet Gonsior, Julius
Falkenberg, Christian
Magino, Silvio
Reusch, Anja
Thiele, Maik
Lehner, Wolfgang
contents Despite achieving state-of-the-art results in nearly all Natural Language Processing applications, fine-tuning Transformer-based language models still requires a significant amount of labeled data to work. A well known technique to reduce the amount of human effort in acquiring a labeled dataset is \textit{Active Learning} (AL): an iterative process in which only the minimal amount of samples is labeled. AL strategies require access to a quantified confidence measure of the model predictions. A common choice is the softmax activation function for the final layer. As the softmax function provides misleading probabilities, this paper compares eight alternatives on seven datasets. Our almost paradoxical finding is that most of the methods are too good at identifying the true most uncertain samples (outliers), and that labeling therefore exclusively outliers results in worse performance. As a heuristic we propose to systematically ignore samples, which results in improvements of various methods compared to the softmax function.
format Preprint
id arxiv_https___arxiv_org_abs_2210_03005
institution arXiv
publishDate 2022
record_format arxiv
spellingShingle To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
Gonsior, Julius
Falkenberg, Christian
Magino, Silvio
Reusch, Anja
Thiele, Maik
Lehner, Wolfgang
Machine Learning
Artificial Intelligence
Computation and Language
Databases
Despite achieving state-of-the-art results in nearly all Natural Language Processing applications, fine-tuning Transformer-based language models still requires a significant amount of labeled data to work. A well known technique to reduce the amount of human effort in acquiring a labeled dataset is \textit{Active Learning} (AL): an iterative process in which only the minimal amount of samples is labeled. AL strategies require access to a quantified confidence measure of the model predictions. A common choice is the softmax activation function for the final layer. As the softmax function provides misleading probabilities, this paper compares eight alternatives on seven datasets. Our almost paradoxical finding is that most of the methods are too good at identifying the true most uncertain samples (outliers), and that labeling therefore exclusively outliers results in worse performance. As a heuristic we propose to systematically ignore samples, which results in improvements of various methods compared to the softmax function.
title To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Databases
url https://arxiv.org/abs/2210.03005