Saved in:
Bibliographic Details
Main Authors: Liu, Wei, Niu, Zhongyu, Gao, Lang, Deng, Zhiying, Wang, Jun, Wang, Haozhao, Li, Ruixuan
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2505.02118
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913977514065920
author Liu, Wei
Niu, Zhongyu
Gao, Lang
Deng, Zhiying
Wang, Jun
Wang, Haozhao
Li, Ruixuan
author_facet Liu, Wei
Niu, Zhongyu
Gao, Lang
Deng, Zhiying
Wang, Jun
Wang, Haozhao
Li, Ruixuan
contents This study investigates the self-rationalization framework constructed with a cooperative game, where a generator initially extracts the most informative segment from raw input, and a subsequent predictor utilizes the selected subset for its input. The generator and predictor are trained collaboratively to maximize prediction accuracy. In this paper, we first uncover a potential caveat: such a cooperative game could unintentionally introduce a sampling bias during rationale extraction. Specifically, the generator might inadvertently create an incorrect correlation between the selected rationale candidate and the label, even when they are semantically unrelated in the original dataset. Subsequently, we elucidate the origins of this bias using both detailed theoretical analysis and empirical evidence. Our findings suggest a direction for inspecting these correlations through attacks, based on which we further introduce an instruction to prevent the predictor from learning the correlations. Through experiments on six text classification datasets and two graph classification datasets using three network architectures (GRUs, BERT, and GCN), we show that our method not only significantly outperforms recent rationalization methods, but also achieves comparable or even better results than a representative LLM (llama3.1-8b-instruct).
format Preprint
id arxiv_https___arxiv_org_abs_2505_02118
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean Datasets
Liu, Wei
Niu, Zhongyu
Gao, Lang
Deng, Zhiying
Wang, Jun
Wang, Haozhao
Li, Ruixuan
Artificial Intelligence
This study investigates the self-rationalization framework constructed with a cooperative game, where a generator initially extracts the most informative segment from raw input, and a subsequent predictor utilizes the selected subset for its input. The generator and predictor are trained collaboratively to maximize prediction accuracy. In this paper, we first uncover a potential caveat: such a cooperative game could unintentionally introduce a sampling bias during rationale extraction. Specifically, the generator might inadvertently create an incorrect correlation between the selected rationale candidate and the label, even when they are semantically unrelated in the original dataset. Subsequently, we elucidate the origins of this bias using both detailed theoretical analysis and empirical evidence. Our findings suggest a direction for inspecting these correlations through attacks, based on which we further introduce an instruction to prevent the predictor from learning the correlations. Through experiments on six text classification datasets and two graph classification datasets using three network architectures (GRUs, BERT, and GCN), we show that our method not only significantly outperforms recent rationalization methods, but also achieves comparable or even better results than a representative LLM (llama3.1-8b-instruct).
title Adversarial Cooperative Rationalization: The Risk of Spurious Correlations in Even Clean Datasets
topic Artificial Intelligence
url https://arxiv.org/abs/2505.02118