Low-Resource Fast Text Classification Based on Intra-Class and Inter-Class Distance Calculation

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Mao, Yanxu, Liu, Peipei, Cui, Tiehan, Liu, Congying, You, Datao
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915062851043328
author Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Liu, Congying
You, Datao
author_facet Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Liu, Congying
You, Datao
contents In recent years, text classification methods based on neural networks and pre-trained models have gained increasing attention and demonstrated excellent performance. However, these methods still have some limitations in practical applications: (1) They typically focus only on the matching similarity between sentences. However, there exists implicit high-value information both within sentences of the same class and across different classes, which is very crucial for classification tasks. (2) Existing methods such as pre-trained language models and graph-based approaches often consume substantial memory for training and text-graph construction. (3) Although some low-resource methods can achieve good performance, they often suffer from excessively long processing times. To address these challenges, we propose a low-resource and fast text classification model called LFTC. Our approach begins by constructing a compressor list for each class to fully mine the regularity information within intra-class data. We then remove redundant information irrelevant to the target classification to reduce processing time. Finally, we compute the similarity distance between text pairs for classification. We evaluate LFTC on 9 publicly available benchmark datasets, and the results demonstrate significant improvements in performance and processing time, especially under limited computational and data resources, highlighting its superior advantages.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09922
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Low-Resource Fast Text Classification Based on Intra-Class and Inter-Class Distance Calculation
Mao, Yanxu
Liu, Peipei
Cui, Tiehan
Liu, Congying
You, Datao
Computation and Language
In recent years, text classification methods based on neural networks and pre-trained models have gained increasing attention and demonstrated excellent performance. However, these methods still have some limitations in practical applications: (1) They typically focus only on the matching similarity between sentences. However, there exists implicit high-value information both within sentences of the same class and across different classes, which is very crucial for classification tasks. (2) Existing methods such as pre-trained language models and graph-based approaches often consume substantial memory for training and text-graph construction. (3) Although some low-resource methods can achieve good performance, they often suffer from excessively long processing times. To address these challenges, we propose a low-resource and fast text classification model called LFTC. Our approach begins by constructing a compressor list for each class to fully mine the regularity information within intra-class data. We then remove redundant information irrelevant to the target classification to reduce processing time. Finally, we compute the similarity distance between text pairs for classification. We evaluate LFTC on 9 publicly available benchmark datasets, and the results demonstrate significant improvements in performance and processing time, especially under limited computational and data resources, highlighting its superior advantages.
title Low-Resource Fast Text Classification Based on Intra-Class and Inter-Class Distance Calculation
topic Computation and Language
url https://arxiv.org/abs/2412.09922