IDoFew: Intermediate Training Using Dual-Clustering in Language Models for Few Labels Text Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alsuhaibani, Abdullah, Zogan, Hamad, Razzak, Imran, Jameel, Shoaib, Xu, Guandong
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916084383219712
author Alsuhaibani, Abdullah
Zogan, Hamad
Razzak, Imran
Jameel, Shoaib
Xu, Guandong
author_facet Alsuhaibani, Abdullah
Zogan, Hamad
Razzak, Imran
Jameel, Shoaib
Xu, Guandong
contents Language models such as Bidirectional Encoder Representations from Transformers (BERT) have been very effective in various Natural Language Processing (NLP) and text mining tasks including text classification. However, some tasks still pose challenges for these models, including text classification with limited labels. This can result in a cold-start problem. Although some approaches have attempted to address this problem through single-stage clustering as an intermediate training step coupled with a pre-trained language model, which generates pseudo-labels to improve classification, these methods are often error-prone due to the limitations of the clustering algorithms. To overcome this, we have developed a novel two-stage intermediate clustering with subsequent fine-tuning that models the pseudo-labels reliably, resulting in reduced prediction errors. The key novelty in our model, IDoFew, is that the two-stage clustering coupled with two different clustering algorithms helps exploit the advantages of the complementary algorithms that reduce the errors in generating reliable pseudo-labels for fine-tuning. Our approach has shown significant improvements compared to strong comparative models.
format Preprint
id arxiv_https___arxiv_org_abs_2401_04025
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle IDoFew: Intermediate Training Using Dual-Clustering in Language Models for Few Labels Text Classification
Alsuhaibani, Abdullah
Zogan, Hamad
Razzak, Imran
Jameel, Shoaib
Xu, Guandong
Computation and Language
Language models such as Bidirectional Encoder Representations from Transformers (BERT) have been very effective in various Natural Language Processing (NLP) and text mining tasks including text classification. However, some tasks still pose challenges for these models, including text classification with limited labels. This can result in a cold-start problem. Although some approaches have attempted to address this problem through single-stage clustering as an intermediate training step coupled with a pre-trained language model, which generates pseudo-labels to improve classification, these methods are often error-prone due to the limitations of the clustering algorithms. To overcome this, we have developed a novel two-stage intermediate clustering with subsequent fine-tuning that models the pseudo-labels reliably, resulting in reduced prediction errors. The key novelty in our model, IDoFew, is that the two-stage clustering coupled with two different clustering algorithms helps exploit the advantages of the complementary algorithms that reduce the errors in generating reliable pseudo-labels for fine-tuning. Our approach has shown significant improvements compared to strong comparative models.
title IDoFew: Intermediate Training Using Dual-Clustering in Language Models for Few Labels Text Classification
topic Computation and Language
url https://arxiv.org/abs/2401.04025