Epistemic Integrity in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ghafouri, Bijean, Mohammadzadeh, Shahrad, Zhou, James, Nair, Pratheeksha, Tian, Jacob-Junqi, Tsujimura, Hikaru, Goel, Mayank, Krishna, Sukanya, Rabbany, Reihaneh, Godbout, Jean-François, Pelrine, Kellin
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908398209990656
author Ghafouri, Bijean
Mohammadzadeh, Shahrad
Zhou, James
Nair, Pratheeksha
Tian, Jacob-Junqi
Tsujimura, Hikaru
Goel, Mayank
Krishna, Sukanya
Rabbany, Reihaneh
Godbout, Jean-François
Pelrine, Kellin
author_facet Ghafouri, Bijean
Mohammadzadeh, Shahrad
Zhou, James
Nair, Pratheeksha
Tian, Jacob-Junqi
Tsujimura, Hikaru
Goel, Mayank
Krishna, Sukanya
Rabbany, Reihaneh
Godbout, Jean-François
Pelrine, Kellin
contents Large language models are increasingly relied upon as sources of information, but their propensity for generating false or misleading statements with high confidence poses risks for users and society. In this paper, we confront the critical problem of epistemic miscalibration $\unicode{x2013}$ where a model's linguistic assertiveness fails to reflect its true internal certainty. We introduce a new human-labeled dataset and a novel method for measuring the linguistic assertiveness of Large Language Models (LLMs) which cuts error rates by over 50% relative to previous benchmarks. Validated across multiple datasets, our method reveals a stark misalignment between how confidently models linguistically present information and their actual accuracy. Further human evaluations confirm the severity of this miscalibration. This evidence underscores the urgent risk of the overstated certainty LLMs hold which may mislead users on a massive scale. Our framework provides a crucial step forward in diagnosing this miscalibration, offering a path towards correcting it and more trustworthy AI across domains.
format Preprint
id arxiv_https___arxiv_org_abs_2411_06528
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Epistemic Integrity in Large Language Models
Ghafouri, Bijean
Mohammadzadeh, Shahrad
Zhou, James
Nair, Pratheeksha
Tian, Jacob-Junqi
Tsujimura, Hikaru
Goel, Mayank
Krishna, Sukanya
Rabbany, Reihaneh
Godbout, Jean-François
Pelrine, Kellin
Computation and Language
Artificial Intelligence
Human-Computer Interaction
Large language models are increasingly relied upon as sources of information, but their propensity for generating false or misleading statements with high confidence poses risks for users and society. In this paper, we confront the critical problem of epistemic miscalibration $\unicode{x2013}$ where a model's linguistic assertiveness fails to reflect its true internal certainty. We introduce a new human-labeled dataset and a novel method for measuring the linguistic assertiveness of Large Language Models (LLMs) which cuts error rates by over 50% relative to previous benchmarks. Validated across multiple datasets, our method reveals a stark misalignment between how confidently models linguistically present information and their actual accuracy. Further human evaluations confirm the severity of this miscalibration. This evidence underscores the urgent risk of the overstated certainty LLMs hold which may mislead users on a massive scale. Our framework provides a crucial step forward in diagnosing this miscalibration, offering a path towards correcting it and more trustworthy AI across domains.
title Epistemic Integrity in Large Language Models
topic Computation and Language
Artificial Intelligence
Human-Computer Interaction
url https://arxiv.org/abs/2411.06528