CBEval: A framework for evaluating and interpreting cognitive biases in LLMs

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shaikh, Ammar, Dandekar, Raj Abhijit, Panat, Sreedath, Dandekar, Rajat
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915049220603904
author Shaikh, Ammar
Dandekar, Raj Abhijit
Panat, Sreedath
Dandekar, Rajat
author_facet Shaikh, Ammar
Dandekar, Raj Abhijit
Panat, Sreedath
Dandekar, Rajat
contents Rapid advancements in Large Language models (LLMs) has significantly enhanced their reasoning capabilities. Despite improved performance on benchmarks, LLMs exhibit notable gaps in their cognitive processes. Additionally, as reflections of human-generated data, these models have the potential to inherit cognitive biases, raising concerns about their reasoning and decision making capabilities. In this paper we present a framework to interpret, understand and provide insights into a host of cognitive biases in LLMs. Conducting our research on frontier language models we're able to elucidate reasoning limitations and biases, and provide reasoning behind these biases by constructing influence graphs that identify phrases and words most responsible for biases manifested in LLMs. We further investigate biases such as round number bias and cognitive bias barrier revealed when noting framing effect in language models.
format Preprint
id arxiv_https___arxiv_org_abs_2412_03605
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
Shaikh, Ammar
Dandekar, Raj Abhijit
Panat, Sreedath
Dandekar, Rajat
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Rapid advancements in Large Language models (LLMs) has significantly enhanced their reasoning capabilities. Despite improved performance on benchmarks, LLMs exhibit notable gaps in their cognitive processes. Additionally, as reflections of human-generated data, these models have the potential to inherit cognitive biases, raising concerns about their reasoning and decision making capabilities. In this paper we present a framework to interpret, understand and provide insights into a host of cognitive biases in LLMs. Conducting our research on frontier language models we're able to elucidate reasoning limitations and biases, and provide reasoning behind these biases by constructing influence graphs that identify phrases and words most responsible for biases manifested in LLMs. We further investigate biases such as round number bias and cognitive bias barrier revealed when noting framing effect in language models.
title CBEval: A framework for evaluating and interpreting cognitive biases in LLMs
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2412.03605