TARDiS : Text Augmentation for Refining Diversity and Separability

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Kim, Kyungmin, Im, SangHun, Kim, GiBaeg, Oh, Heung-Seon
Format: Preprint
Publié: 2025
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866909448619950080
author Kim, Kyungmin
Im, SangHun
Kim, GiBaeg
Oh, Heung-Seon
author_facet Kim, Kyungmin
Im, SangHun
Kim, GiBaeg
Oh, Heung-Seon
contents Text augmentation (TA) is a critical technique for text classification, especially in few-shot settings. This paper introduces a novel LLM-based TA method, TARDiS, to address challenges inherent in the generation and alignment stages of two-stage TA methods. For the generation stage, we propose two generation processes, SEG and CEG, incorporating multiple class-specific prompts to enhance diversity and separability. For the alignment stage, we introduce a class adaptation (CA) method to ensure that generated examples align with their target classes through verification and modification. Experimental results demonstrate TARDiS's effectiveness, outperforming state-of-the-art LLM-based TA methods in various few-shot text classification tasks. An in-depth analysis confirms the detailed behaviors at each stage.
format Preprint
id arxiv_https___arxiv_org_abs_2501_02739
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle TARDiS : Text Augmentation for Refining Diversity and Separability
Kim, Kyungmin
Im, SangHun
Kim, GiBaeg
Oh, Heung-Seon
Computation and Language
Artificial Intelligence
Machine Learning
Text augmentation (TA) is a critical technique for text classification, especially in few-shot settings. This paper introduces a novel LLM-based TA method, TARDiS, to address challenges inherent in the generation and alignment stages of two-stage TA methods. For the generation stage, we propose two generation processes, SEG and CEG, incorporating multiple class-specific prompts to enhance diversity and separability. For the alignment stage, we introduce a class adaptation (CA) method to ensure that generated examples align with their target classes through verification and modification. Experimental results demonstrate TARDiS's effectiveness, outperforming state-of-the-art LLM-based TA methods in various few-shot text classification tasks. An in-depth analysis confirms the detailed behaviors at each stage.
title TARDiS : Text Augmentation for Refining Diversity and Separability
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2501.02739