KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kim, Seorin, Lee, Dongyoung, Lee, Jaejin
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913960877359104
author Kim, Seorin
Lee, Dongyoung
Lee, Jaejin
author_facet Kim, Seorin
Lee, Dongyoung
Lee, Jaejin
contents Large language models (LLMs) often exhibit societal biases in their outputs, prompting ethical concerns regarding fairness and harm. In this work, we propose KLAAD (KL-Attention Alignment Debiasing), an attention-based debiasing framework that implicitly aligns attention distributions between stereotypical and anti-stereotypical sentence pairs without directly modifying model weights. KLAAD introduces a composite training objective combining Cross-Entropy, KL divergence, and Triplet losses, guiding the model to consistently attend across biased and unbiased contexts while preserving fluency and coherence. Experimental evaluation of KLAAD demonstrates improved bias mitigation on both the BBQ and BOLD benchmarks, with minimal impact on language modeling quality. The results indicate that attention-level alignment offers a principled solution for mitigating bias in generative language models.
format Preprint
id arxiv_https___arxiv_org_abs_2507_19962
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
Kim, Seorin
Lee, Dongyoung
Lee, Jaejin
Computation and Language
Large language models (LLMs) often exhibit societal biases in their outputs, prompting ethical concerns regarding fairness and harm. In this work, we propose KLAAD (KL-Attention Alignment Debiasing), an attention-based debiasing framework that implicitly aligns attention distributions between stereotypical and anti-stereotypical sentence pairs without directly modifying model weights. KLAAD introduces a composite training objective combining Cross-Entropy, KL divergence, and Triplet losses, guiding the model to consistently attend across biased and unbiased contexts while preserving fluency and coherence. Experimental evaluation of KLAAD demonstrates improved bias mitigation on both the BBQ and BOLD benchmarks, with minimal impact on language modeling quality. The results indicate that attention-level alignment offers a principled solution for mitigating bias in generative language models.
title KLAAD: Refining Attention Mechanisms to Reduce Societal Bias in Generative Language Models
topic Computation and Language
url https://arxiv.org/abs/2507.19962