Explore Spurious Correlations at the Concept Level in Language Models for Text Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhou, Yuhang, Xu, Paiheng, Liu, Xiaoyu, An, Bang, Ai, Wei, Huang, Furong
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929386929782784
author Zhou, Yuhang
Xu, Paiheng
Liu, Xiaoyu
An, Bang
Ai, Wei
Huang, Furong
author_facet Zhou, Yuhang
Xu, Paiheng
Liu, Xiaoyu
An, Bang
Ai, Wei
Huang, Furong
contents Language models (LMs) have achieved notable success in numerous NLP tasks, employing both fine-tuning and in-context learning (ICL) methods. While language models demonstrate exceptional performance, they face robustness challenges due to spurious correlations arising from imbalanced label distributions in training data or ICL exemplars. Previous research has primarily concentrated on word, phrase, and syntax features, neglecting the concept level, often due to the absence of concept labels and difficulty in identifying conceptual content in input texts. This paper introduces two main contributions. First, we employ ChatGPT to assign concept labels to texts, assessing concept bias in models during fine-tuning or ICL on test data. We find that LMs, when encountering spurious correlations between a concept and a label in training or prompts, resort to shortcuts for predictions. Second, we introduce a data rebalancing technique that incorporates ChatGPT-generated counterfactual data, thereby balancing label distribution and mitigating spurious correlations. Our method's efficacy, surpassing traditional token removal approaches, is validated through extensive testing.
format Preprint
id arxiv_https___arxiv_org_abs_2311_08648
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Explore Spurious Correlations at the Concept Level in Language Models for Text Classification
Zhou, Yuhang
Xu, Paiheng
Liu, Xiaoyu
An, Bang
Ai, Wei
Huang, Furong
Computation and Language
Artificial Intelligence
Language models (LMs) have achieved notable success in numerous NLP tasks, employing both fine-tuning and in-context learning (ICL) methods. While language models demonstrate exceptional performance, they face robustness challenges due to spurious correlations arising from imbalanced label distributions in training data or ICL exemplars. Previous research has primarily concentrated on word, phrase, and syntax features, neglecting the concept level, often due to the absence of concept labels and difficulty in identifying conceptual content in input texts. This paper introduces two main contributions. First, we employ ChatGPT to assign concept labels to texts, assessing concept bias in models during fine-tuning or ICL on test data. We find that LMs, when encountering spurious correlations between a concept and a label in training or prompts, resort to shortcuts for predictions. Second, we introduce a data rebalancing technique that incorporates ChatGPT-generated counterfactual data, thereby balancing label distribution and mitigating spurious correlations. Our method's efficacy, surpassing traditional token removal approaches, is validated through extensive testing.
title Explore Spurious Correlations at the Concept Level in Language Models for Text Classification
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2311.08648