Information Gain-Guided Causal Intervention for Autonomous Debiasing Large Language Models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Sun, Zhouhao, Ding, Xiao, Du, Li, Xu, Yunpeng, Ma, Yixuan, Zhao, Yang, Qin, Bing, Liu, Ting
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916760109711360
author Sun, Zhouhao
Ding, Xiao
Du, Li
Xu, Yunpeng
Ma, Yixuan
Zhao, Yang
Qin, Bing
Liu, Ting
author_facet Sun, Zhouhao
Ding, Xiao
Du, Li
Xu, Yunpeng
Ma, Yixuan
Zhao, Yang
Qin, Bing
Liu, Ting
contents Despite significant progress, recent studies indicate that current large language models (LLMs) may still capture dataset biases and utilize them during inference, leading to the poor generalizability of LLMs. However, due to the diversity of dataset biases and the insufficient nature of bias suppression based on in-context learning, the effectiveness of previous prior knowledge-based debiasing methods and in-context learning based automatic debiasing methods is limited. To address these challenges, we explore the combination of causal mechanisms with information theory and propose an information gain-guided causal intervention debiasing (ICD) framework. To eliminate biases within the instruction-tuning dataset, it is essential to ensure that these biases do not provide any additional information to predict the answers, i.e., the information gain of these biases for predicting the answers needs to be 0. Under this guidance, this framework utilizes a causal intervention-based data rewriting method to automatically and autonomously balance the distribution of instruction-tuning dataset for reducing the information gain. Subsequently, it employs a standard supervised fine-tuning process to train LLMs on the debiased dataset. Experimental results show that ICD can effectively debias LLM to improve its generalizability across different tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2504_12898
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Information Gain-Guided Causal Intervention for Autonomous Debiasing Large Language Models
Sun, Zhouhao
Ding, Xiao
Du, Li
Xu, Yunpeng
Ma, Yixuan
Zhao, Yang
Qin, Bing
Liu, Ting
Computation and Language
Artificial Intelligence
Despite significant progress, recent studies indicate that current large language models (LLMs) may still capture dataset biases and utilize them during inference, leading to the poor generalizability of LLMs. However, due to the diversity of dataset biases and the insufficient nature of bias suppression based on in-context learning, the effectiveness of previous prior knowledge-based debiasing methods and in-context learning based automatic debiasing methods is limited. To address these challenges, we explore the combination of causal mechanisms with information theory and propose an information gain-guided causal intervention debiasing (ICD) framework. To eliminate biases within the instruction-tuning dataset, it is essential to ensure that these biases do not provide any additional information to predict the answers, i.e., the information gain of these biases for predicting the answers needs to be 0. Under this guidance, this framework utilizes a causal intervention-based data rewriting method to automatically and autonomously balance the distribution of instruction-tuning dataset for reducing the information gain. Subsequently, it employs a standard supervised fine-tuning process to train LLMs on the debiased dataset. Experimental results show that ICD can effectively debias LLM to improve its generalizability across different tasks.
title Information Gain-Guided Causal Intervention for Autonomous Debiasing Large Language Models
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2504.12898