From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Kabir, Mohsinul, Tahsin, Tasfia, Ananiadou, Sophia
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909900054986752
author Kabir, Mohsinul
Tahsin, Tasfia
Ananiadou, Sophia
author_facet Kabir, Mohsinul
Tahsin, Tasfia
Ananiadou, Sophia
contents Current research on bias in language models (LMs) predominantly focuses on data quality, with significantly less attention paid to model architecture and temporal influences of data. Even more critically, few studies systematically investigate the origins of bias. We propose a methodology grounded in comparative behavioral theory to interpret the complex interaction between training data and model architecture in bias propagation during language modeling. Building on recent work that relates transformers to n-gram LMs, we evaluate how data, model design choices, and temporal dynamics affect bias propagation. Our findings reveal that: (1) n-gram LMs are highly sensitive to context window size in bias propagation, while transformers demonstrate architectural robustness; (2) the temporal provenance of training data significantly affects bias; and (3) different model architectures respond differentially to controlled bias injection, with certain biases (e.g. sexual orientation) being disproportionately amplified. As language models become ubiquitous, our findings highlight the need for a holistic approach -- tracing bias to its origins across both data and model dimensions, not just symptoms, to mitigate harm.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12381
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
Kabir, Mohsinul
Tahsin, Tasfia
Ananiadou, Sophia
Computation and Language
Artificial Intelligence
Current research on bias in language models (LMs) predominantly focuses on data quality, with significantly less attention paid to model architecture and temporal influences of data. Even more critically, few studies systematically investigate the origins of bias. We propose a methodology grounded in comparative behavioral theory to interpret the complex interaction between training data and model architecture in bias propagation during language modeling. Building on recent work that relates transformers to n-gram LMs, we evaluate how data, model design choices, and temporal dynamics affect bias propagation. Our findings reveal that: (1) n-gram LMs are highly sensitive to context window size in bias propagation, while transformers demonstrate architectural robustness; (2) the temporal provenance of training data significantly affects bias; and (3) different model architectures respond differentially to controlled bias injection, with certain biases (e.g. sexual orientation) being disproportionately amplified. As language models become ubiquitous, our findings highlight the need for a holistic approach -- tracing bias to its origins across both data and model dimensions, not just symptoms, to mitigate harm.
title From n-gram to Attention: How Model Architectures Learn and Propagate Bias in Language Modeling
topic Computation and Language
Artificial Intelligence
url https://arxiv.org/abs/2505.12381