Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | , , , , |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
| _version_ | 1866909199204614144 |
|---|---|
| author | Jenny, David F. Billeter, Yann Sachan, Mrinmaya Schölkopf, Bernhard Jin, Zhijing |
| author_facet | Jenny, David F. Billeter, Yann Sachan, Mrinmaya Schölkopf, Bernhard Jin, Zhijing |
| contents | The rapid advancement of Large Language Models (LLMs) has sparked intense debate regarding the prevalence of bias in these models and its mitigation. Yet, as exemplified by both results on debiasing methods in the literature and reports of alignment-related defects from the wider community, bias remains a poorly understood topic despite its practical relevance. To enhance the understanding of the internal causes of bias, we analyse LLM bias through the lens of causal fairness analysis, which enables us to both comprehend the origins of bias and reason about its downstream consequences and mitigation. To operationalize this framework, we propose a prompt-based method for the extraction of confounding and mediating attributes which contribute to the LLM decision process. By applying Activity Dependency Networks (ADNs), we then analyse how these attributes influence an LLM's decision process. We apply our method to LLM ratings of argument quality in political debates. We find that the observed disparate treatment can at least in part be attributed to confounding and mitigating attributes and model misalignment, and discuss the consequences of our findings for human-AI alignment and bias mitigation. Our code and data are at https://github.com/david-jenny/LLM-Political-Study. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2311_08605 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis Jenny, David F. Billeter, Yann Sachan, Mrinmaya Schölkopf, Bernhard Jin, Zhijing Computation and Language Artificial Intelligence Computers and Society Social and Information Networks The rapid advancement of Large Language Models (LLMs) has sparked intense debate regarding the prevalence of bias in these models and its mitigation. Yet, as exemplified by both results on debiasing methods in the literature and reports of alignment-related defects from the wider community, bias remains a poorly understood topic despite its practical relevance. To enhance the understanding of the internal causes of bias, we analyse LLM bias through the lens of causal fairness analysis, which enables us to both comprehend the origins of bias and reason about its downstream consequences and mitigation. To operationalize this framework, we propose a prompt-based method for the extraction of confounding and mediating attributes which contribute to the LLM decision process. By applying Activity Dependency Networks (ADNs), we then analyse how these attributes influence an LLM's decision process. We apply our method to LLM ratings of argument quality in political debates. We find that the observed disparate treatment can at least in part be attributed to confounding and mitigating attributes and model misalignment, and discuss the consequences of our findings for human-AI alignment and bias mitigation. Our code and data are at https://github.com/david-jenny/LLM-Political-Study. |
| title | Exploring the Jungle of Bias: Political Bias Attribution in Language Models via Dependency Analysis |
| topic | Computation and Language Artificial Intelligence Computers and Society Social and Information Networks |
| url | https://arxiv.org/abs/2311.08605 |