Fair In-Context Learning via Latent Concept Variables

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bhaila, Karuna, Van, Minh-Hao, Edemacu, Kennedy, Zhao, Chen, Chen, Feng, Wu, Xintao
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917080405639168
author Bhaila, Karuna
Van, Minh-Hao
Edemacu, Kennedy
Zhao, Chen
Chen, Feng
Wu, Xintao
author_facet Bhaila, Karuna
Van, Minh-Hao
Edemacu, Kennedy
Zhao, Chen
Chen, Feng
Wu, Xintao
contents The emerging in-context learning (ICL) ability of large language models (LLMs) has prompted their use for predictive tasks in various domains with different data types, including tabular data, facilitated by serialization methods. However, with increasing applications in high-stakes domains, it has been shown that LLMs can inherit social bias and discrimination from their pre-training data. In this work, we investigate inherent bias in LLMs during in-context learning with tabular data. We focus on an optimal demonstration selection approach that utilizes latent concept variables for resource-efficient task adaptation. We design data augmentation strategies that reduce the correlation between predictive outcomes and sensitive variables, helping promote fairness during latent concept learning. We utilize the learned concept to select demonstrations and obtain fair predictions. The latent concept variables are learned using a smaller internal LLM and generalized to larger external LLMs. We empirically verify that the fair latent variable approach improves fairness results on tabular datasets compared to multiple heuristic demonstration selection methods.
format Preprint
id arxiv_https___arxiv_org_abs_2411_02671
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Fair In-Context Learning via Latent Concept Variables
Bhaila, Karuna
Van, Minh-Hao
Edemacu, Kennedy
Zhao, Chen
Chen, Feng
Wu, Xintao
Machine Learning
Artificial Intelligence
Computation and Language
The emerging in-context learning (ICL) ability of large language models (LLMs) has prompted their use for predictive tasks in various domains with different data types, including tabular data, facilitated by serialization methods. However, with increasing applications in high-stakes domains, it has been shown that LLMs can inherit social bias and discrimination from their pre-training data. In this work, we investigate inherent bias in LLMs during in-context learning with tabular data. We focus on an optimal demonstration selection approach that utilizes latent concept variables for resource-efficient task adaptation. We design data augmentation strategies that reduce the correlation between predictive outcomes and sensitive variables, helping promote fairness during latent concept learning. We utilize the learned concept to select demonstrations and obtain fair predictions. The latent concept variables are learned using a smaller internal LLM and generalized to larger external LLMs. We empirically verify that the fair latent variable approach improves fairness results on tabular datasets compared to multiple heuristic demonstration selection methods.
title Fair In-Context Learning via Latent Concept Variables
topic Machine Learning
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2411.02671