Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Park, Eunkyu, Deng, Wesley Hanwen, Varadarajan, Vasudha, Yan, Mingxi, Kim, Gunhee, Sap, Maarten, Eslami, Motahhare
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911275878973440
author Park, Eunkyu
Deng, Wesley Hanwen
Varadarajan, Vasudha
Yan, Mingxi
Kim, Gunhee
Sap, Maarten
Eslami, Motahhare
author_facet Park, Eunkyu
Deng, Wesley Hanwen
Varadarajan, Vasudha
Yan, Mingxi
Kim, Gunhee
Sap, Maarten
Eslami, Motahhare
contents Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of-Thought (CoT) explanations in multimodal moral scenarios by systematically perturbing reasoning chains and manipulating delivery tones. Specifically, we analyze reasoning errors in vision language models (VLMs) and how they impact user trust and the ability to detect errors. Our findings reveal two key effects: (1) users often equate trust with outcome agreement, sustaining reliance even when reasoning is flawed, and (2) the confident tone suppresses error detection while maintaining reliance, showing that delivery styles can override correctness. These results highlight how CoT explanations can simultaneously clarify and mislead, underscoring the need for NLP systems to provide explanations that encourage scrutiny and critical thinking rather than blind trust. All code will be released publicly.
format Preprint
id arxiv_https___arxiv_org_abs_2511_12001
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
Park, Eunkyu
Deng, Wesley Hanwen
Varadarajan, Vasudha
Yan, Mingxi
Kim, Gunhee
Sap, Maarten
Eslami, Motahhare
Computation and Language
Human-Computer Interaction
Explanations are often promoted as tools for transparency, but they can also foster confirmation bias; users may assume reasoning is correct whenever outputs appear acceptable. We study this double-edged role of Chain-of-Thought (CoT) explanations in multimodal moral scenarios by systematically perturbing reasoning chains and manipulating delivery tones. Specifically, we analyze reasoning errors in vision language models (VLMs) and how they impact user trust and the ability to detect errors. Our findings reveal two key effects: (1) users often equate trust with outcome agreement, sustaining reliance even when reasoning is flawed, and (2) the confident tone suppresses error detection while maintaining reliance, showing that delivery styles can override correctness. These results highlight how CoT explanations can simultaneously clarify and mislead, underscoring the need for NLP systems to provide explanations that encourage scrutiny and critical thinking rather than blind trust. All code will be released publicly.
title Critical or Compliant? The Double-Edged Sword of Reasoning in Chain-of-Thought Explanations
topic Computation and Language
Human-Computer Interaction
url https://arxiv.org/abs/2511.12001