Towards Universal Debiasing for Language Models-based Tabular Data Generation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Tianchun, Liu, Tianci, Wang, Xingchen, Wei, Rongzhe, Li, Pan, Su, Lu, Gao, Jing
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912595596804096
author Li, Tianchun
Liu, Tianci
Wang, Xingchen
Wei, Rongzhe
Li, Pan
Su, Lu
Gao, Jing
author_facet Li, Tianchun
Liu, Tianci
Wang, Xingchen
Wei, Rongzhe
Li, Pan
Su, Lu
Gao, Jing
contents Large language models (LLMs) have achieved promising results in tabular data generation. However, inherent historical biases in tabular datasets often cause LLMs to exacerbate fairness issues, particularly when multiple advantaged and protected features are involved. In this work, we introduce a universal debiasing framework that minimizes group-level dependencies by simultaneously reducing the mutual information between advantaged and protected attributes. By leveraging the autoregressive structure and analytic sampling distributions of LLM-based tabular data generators, our approach efficiently computes mutual information, reducing the need for cumbersome numerical estimations. Building on this foundation, we propose two complementary methods: a direct preference optimization (DPO)-based strategy, namely UDF-DPO, that integrates seamlessly with existing models, and a targeted debiasing technique, namely UDF-MIX, that achieves debiasing without tuning the parameters of LLMs. Extensive experiments demonstrate that our framework effectively balances fairness and utility, offering a scalable and practical solution for debiasing in high-stakes applications.
format Preprint
id arxiv_https___arxiv_org_abs_2509_16475
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Universal Debiasing for Language Models-based Tabular Data Generation
Li, Tianchun
Liu, Tianci
Wang, Xingchen
Wei, Rongzhe
Li, Pan
Su, Lu
Gao, Jing
Machine Learning
Computation and Language
Large language models (LLMs) have achieved promising results in tabular data generation. However, inherent historical biases in tabular datasets often cause LLMs to exacerbate fairness issues, particularly when multiple advantaged and protected features are involved. In this work, we introduce a universal debiasing framework that minimizes group-level dependencies by simultaneously reducing the mutual information between advantaged and protected attributes. By leveraging the autoregressive structure and analytic sampling distributions of LLM-based tabular data generators, our approach efficiently computes mutual information, reducing the need for cumbersome numerical estimations. Building on this foundation, we propose two complementary methods: a direct preference optimization (DPO)-based strategy, namely UDF-DPO, that integrates seamlessly with existing models, and a targeted debiasing technique, namely UDF-MIX, that achieves debiasing without tuning the parameters of LLMs. Extensive experiments demonstrate that our framework effectively balances fairness and utility, offering a scalable and practical solution for debiasing in high-stakes applications.
title Towards Universal Debiasing for Language Models-based Tabular Data Generation
topic Machine Learning
Computation and Language
url https://arxiv.org/abs/2509.16475