Thinking Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Furniturewala, Shaz, Jandial, Surgan, Java, Abhinav, Banerjee, Pragyan, Shahid, Simra, Bhatia, Sumit, Jaidka, Kokil
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914799782199296
author Furniturewala, Shaz
Jandial, Surgan
Java, Abhinav
Banerjee, Pragyan
Shahid, Simra
Bhatia, Sumit
Jaidka, Kokil
author_facet Furniturewala, Shaz
Jandial, Surgan
Java, Abhinav
Banerjee, Pragyan
Shahid, Simra
Bhatia, Sumit
Jaidka, Kokil
contents Existing debiasing techniques are typically training-based or require access to the model's internals and output distributions, so they are inaccessible to end-users looking to adapt LLM outputs for their particular needs. In this study, we examine whether structured prompting techniques can offer opportunities for fair text generation. We evaluate a comprehensive end-user-focused iterative framework of debiasing that applies System 2 thinking processes for prompts to induce logical, reflective, and critical text generation, with single, multi-step, instruction, and role-based variants. By systematically evaluating many LLMs across many datasets and different prompting strategies, we show that the more complex System 2-based Implicative Prompts significantly improve over other techniques demonstrating lower mean bias in the outputs with competitive performance on the downstream tasks. Our work offers research directions for the design and the potential of end-user-focused evaluative frameworks for LLM use.
format Preprint
id arxiv_https___arxiv_org_abs_2405_10431
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Thinking Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models
Furniturewala, Shaz
Jandial, Surgan
Java, Abhinav
Banerjee, Pragyan
Shahid, Simra
Bhatia, Sumit
Jaidka, Kokil
Computation and Language
Existing debiasing techniques are typically training-based or require access to the model's internals and output distributions, so they are inaccessible to end-users looking to adapt LLM outputs for their particular needs. In this study, we examine whether structured prompting techniques can offer opportunities for fair text generation. We evaluate a comprehensive end-user-focused iterative framework of debiasing that applies System 2 thinking processes for prompts to induce logical, reflective, and critical text generation, with single, multi-step, instruction, and role-based variants. By systematically evaluating many LLMs across many datasets and different prompting strategies, we show that the more complex System 2-based Implicative Prompts significantly improve over other techniques demonstrating lower mean bias in the outputs with competitive performance on the downstream tasks. Our work offers research directions for the design and the potential of end-user-focused evaluative frameworks for LLM use.
title Thinking Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models
topic Computation and Language
url https://arxiv.org/abs/2405.10431