Revisit Few-shot Intent Classification with PLMs: Direct Fine-tuning vs. Continual Pre-training

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Zhang, Haode, Liang, Haowen, Zhan, Liming, Lam, Albert Y. S., Wu, Xiao-Ming
Format: Preprint
Publié: 2023
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866914948876075008
author Zhang, Haode
Liang, Haowen
Zhan, Liming
Lam, Albert Y. S.
Wu, Xiao-Ming
author_facet Zhang, Haode
Liang, Haowen
Zhan, Liming
Lam, Albert Y. S.
Wu, Xiao-Ming
contents We consider the task of few-shot intent detection, which involves training a deep learning model to classify utterances based on their underlying intents using only a small amount of labeled data. The current approach to address this problem is through continual pre-training, i.e., fine-tuning pre-trained language models (PLMs) on external resources (e.g., conversational corpora, public intent detection datasets, or natural language understanding datasets) before using them as utterance encoders for training an intent classifier. In this paper, we show that continual pre-training may not be essential, since the overfitting problem of PLMs on this task may not be as serious as expected. Specifically, we find that directly fine-tuning PLMs on only a handful of labeled examples already yields decent results compared to methods that employ continual pre-training, and the performance gap diminishes rapidly as the number of labeled data increases. To maximize the utilization of the limited available data, we propose a context augmentation method and leverage sequential self-distillation to boost performance. Comprehensive experiments on real-world benchmarks show that given only two or more labeled samples per class, direct fine-tuning outperforms many strong baselines that utilize external data sources for continual pre-training. The code can be found at https://github.com/hdzhang-code/DFTPlus.
format Preprint
id arxiv_https___arxiv_org_abs_2306_05278
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Revisit Few-shot Intent Classification with PLMs: Direct Fine-tuning vs. Continual Pre-training
Zhang, Haode
Liang, Haowen
Zhan, Liming
Lam, Albert Y. S.
Wu, Xiao-Ming
Computation and Language
We consider the task of few-shot intent detection, which involves training a deep learning model to classify utterances based on their underlying intents using only a small amount of labeled data. The current approach to address this problem is through continual pre-training, i.e., fine-tuning pre-trained language models (PLMs) on external resources (e.g., conversational corpora, public intent detection datasets, or natural language understanding datasets) before using them as utterance encoders for training an intent classifier. In this paper, we show that continual pre-training may not be essential, since the overfitting problem of PLMs on this task may not be as serious as expected. Specifically, we find that directly fine-tuning PLMs on only a handful of labeled examples already yields decent results compared to methods that employ continual pre-training, and the performance gap diminishes rapidly as the number of labeled data increases. To maximize the utilization of the limited available data, we propose a context augmentation method and leverage sequential self-distillation to boost performance. Comprehensive experiments on real-world benchmarks show that given only two or more labeled samples per class, direct fine-tuning outperforms many strong baselines that utilize external data sources for continual pre-training. The code can be found at https://github.com/hdzhang-code/DFTPlus.
title Revisit Few-shot Intent Classification with PLMs: Direct Fine-tuning vs. Continual Pre-training
topic Computation and Language
url https://arxiv.org/abs/2306.05278