Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Itzhak, Itay, Stanovsky, Gabriel, Rosenfeld, Nir, Belinkov, Yonatan
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911819747033088
author Itzhak, Itay
Stanovsky, Gabriel
Rosenfeld, Nir
Belinkov, Yonatan
author_facet Itzhak, Itay
Stanovsky, Gabriel
Rosenfeld, Nir
Belinkov, Yonatan
contents Recent studies show that instruction tuning (IT) and reinforcement learning from human feedback (RLHF) improve the abilities of large language models (LMs) dramatically. While these tuning methods can help align models with human objectives and generate high-quality text, not much is known about their potential adverse effects. In this work, we investigate the effect of IT and RLHF on decision making and reasoning in LMs, focusing on three cognitive biases - the decoy effect, the certainty effect, and the belief bias - all of which are known to influence human decision-making and reasoning. Our findings highlight the presence of these biases in various models from the GPT-3, Mistral, and T5 families. Notably, we find a stronger presence of biases in models that have undergone instruction tuning, such as Flan-T5, Mistral-Instruct, GPT3.5, and GPT4. Our work constitutes a step toward comprehending cognitive biases in instruction-tuned LMs, which is crucial for the development of more reliable and unbiased language models.
format Preprint
id arxiv_https___arxiv_org_abs_2308_00225
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
Itzhak, Itay
Stanovsky, Gabriel
Rosenfeld, Nir
Belinkov, Yonatan
Artificial Intelligence
Computers and Society
Machine Learning
Recent studies show that instruction tuning (IT) and reinforcement learning from human feedback (RLHF) improve the abilities of large language models (LMs) dramatically. While these tuning methods can help align models with human objectives and generate high-quality text, not much is known about their potential adverse effects. In this work, we investigate the effect of IT and RLHF on decision making and reasoning in LMs, focusing on three cognitive biases - the decoy effect, the certainty effect, and the belief bias - all of which are known to influence human decision-making and reasoning. Our findings highlight the presence of these biases in various models from the GPT-3, Mistral, and T5 families. Notably, we find a stronger presence of biases in models that have undergone instruction tuning, such as Flan-T5, Mistral-Instruct, GPT3.5, and GPT4. Our work constitutes a step toward comprehending cognitive biases in instruction-tuned LMs, which is crucial for the development of more reliable and unbiased language models.
title Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
topic Artificial Intelligence
Computers and Society
Machine Learning
url https://arxiv.org/abs/2308.00225