To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bo, Jessica Y., Wan, Sophia, Anderson, Ashton
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912674442379264
author Bo, Jessica Y.
Wan, Sophia
Anderson, Ashton
author_facet Bo, Jessica Y.
Wan, Sophia
Anderson, Ashton
contents As Large Language Models become integral to decision-making, optimism about their power is tempered with concern over their errors. Users may over-rely on LLM advice that is confidently stated but wrong, or under-rely due to mistrust. Reliance interventions have been developed to help users of LLMs, but they lack rigorous evaluation for appropriate reliance. We benchmark the performance of three relevant interventions by conducting a randomized online experiment with 400 participants attempting two challenging tasks: LSAT logical reasoning and image-based numerical estimation. For each question, participants first answered independently, then received LLM advice modified by one of three reliance interventions and answered the question again. Our findings indicate that while interventions reduce over-reliance, they generally fail to improve appropriate reliance. Furthermore, people became more confident after making wrong reliance decisions in certain contexts, demonstrating poor calibration. Based on our findings, we discuss implications for designing effective reliance interventions in human-LLM collaboration.
format Preprint
id arxiv_https___arxiv_org_abs_2412_15584
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models
Bo, Jessica Y.
Wan, Sophia
Anderson, Ashton
Human-Computer Interaction
As Large Language Models become integral to decision-making, optimism about their power is tempered with concern over their errors. Users may over-rely on LLM advice that is confidently stated but wrong, or under-rely due to mistrust. Reliance interventions have been developed to help users of LLMs, but they lack rigorous evaluation for appropriate reliance. We benchmark the performance of three relevant interventions by conducting a randomized online experiment with 400 participants attempting two challenging tasks: LSAT logical reasoning and image-based numerical estimation. For each question, participants first answered independently, then received LLM advice modified by one of three reliance interventions and answered the question again. Our findings indicate that while interventions reduce over-reliance, they generally fail to improve appropriate reliance. Furthermore, people became more confident after making wrong reliance decisions in certain contexts, demonstrating poor calibration. Based on our findings, we discuss implications for designing effective reliance interventions in human-LLM collaboration.
title To Rely or Not to Rely? Evaluating Interventions for Appropriate Reliance on Large Language Models
topic Human-Computer Interaction
url https://arxiv.org/abs/2412.15584