AXOLOTL: Fairness through Assisted Self-Debiasing of Large Language Model Outputs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ebrahimi, Sana, Chen, Kaiwen, Asudeh, Abolfazl, Das, Gautam, Koudas, Nick
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929260041601024
author Ebrahimi, Sana
Chen, Kaiwen
Asudeh, Abolfazl
Das, Gautam
Koudas, Nick
author_facet Ebrahimi, Sana
Chen, Kaiwen
Asudeh, Abolfazl
Das, Gautam
Koudas, Nick
contents Pre-trained Large Language Models (LLMs) have significantly advanced natural language processing capabilities but are susceptible to biases present in their training data, leading to unfair outcomes in various applications. While numerous strategies have been proposed to mitigate bias, they often require extensive computational resources and may compromise model performance. In this work, we introduce AXOLOTL, a novel post-processing framework, which operates agnostically across tasks and models, leveraging public APIs to interact with LLMs without direct access to internal parameters. Through a three-step process resembling zero-shot learning, AXOLOTL identifies biases, proposes resolutions, and guides the model to self-debias its outputs. This approach minimizes computational costs and preserves model performance, making AXOLOTL a promising tool for debiasing LLM outputs with broad applicability and ease of use.
format Preprint
id arxiv_https___arxiv_org_abs_2403_00198
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle AXOLOTL: Fairness through Assisted Self-Debiasing of Large Language Model Outputs
Ebrahimi, Sana
Chen, Kaiwen
Asudeh, Abolfazl
Das, Gautam
Koudas, Nick
Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
Pre-trained Large Language Models (LLMs) have significantly advanced natural language processing capabilities but are susceptible to biases present in their training data, leading to unfair outcomes in various applications. While numerous strategies have been proposed to mitigate bias, they often require extensive computational resources and may compromise model performance. In this work, we introduce AXOLOTL, a novel post-processing framework, which operates agnostically across tasks and models, leveraging public APIs to interact with LLMs without direct access to internal parameters. Through a three-step process resembling zero-shot learning, AXOLOTL identifies biases, proposes resolutions, and guides the model to self-debias its outputs. This approach minimizes computational costs and preserves model performance, making AXOLOTL a promising tool for debiasing LLM outputs with broad applicability and ease of use.
title AXOLOTL: Fairness through Assisted Self-Debiasing of Large Language Model Outputs
topic Computation and Language
Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2403.00198