Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
Fuente:
arXiv
Salvato in:
| Autori principali: | Cheng, Myra, Hawkins, Robert D., Jurafsky, Dan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HumT DumT: Measuring and controlling human-like language in LLMs
di: Cheng, Myra, et al.
Pubblicazione: (2025)
di: Cheng, Myra, et al.
Pubblicazione: (2025)
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
di: Cheng, Myra, et al.
Pubblicazione: (2024)
di: Cheng, Myra, et al.
Pubblicazione: (2024)
ELEPHANT: Measuring and understanding social sycophancy in LLMs
di: Cheng, Myra, et al.
Pubblicazione: (2025)
di: Cheng, Myra, et al.
Pubblicazione: (2025)
Verbalizing LLMs' assumptions to explain and control sycophancy
di: Cheng, Myra, et al.
Pubblicazione: (2026)
di: Cheng, Myra, et al.
Pubblicazione: (2026)
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
di: Suzgun, Mirac, et al.
Pubblicazione: (2024)
di: Suzgun, Mirac, et al.
Pubblicazione: (2024)
Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations
di: Yu, Sunny, et al.
Pubblicazione: (2025)
di: Yu, Sunny, et al.
Pubblicazione: (2025)
When Correct Beliefs Collapse: Epistemic Resilience of LLMs under Clinical Pressure
di: Xiao, Boyu, et al.
Pubblicazione: (2026)
di: Xiao, Boyu, et al.
Pubblicazione: (2026)
Dialect prejudice predicts AI decisions about people's character, employability, and criminality
di: Hofmann, Valentin, et al.
Pubblicazione: (2024)
di: Hofmann, Valentin, et al.
Pubblicazione: (2024)
What can large language models do for sustainable food?
di: Thomas, Anna T., et al.
Pubblicazione: (2025)
di: Thomas, Anna T., et al.
Pubblicazione: (2025)
When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis
di: Dietrich, Juergen
Pubblicazione: (2026)
di: Dietrich, Juergen
Pubblicazione: (2026)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
di: Mohamadi, Alireza, et al.
Pubblicazione: (2025)
di: Mohamadi, Alireza, et al.
Pubblicazione: (2025)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
di: Li, Jing-Jing, et al.
Pubblicazione: (2026)
Sycophantic AI Decreases Prosocial Intentions and Promotes Dependence
di: Cheng, Myra, et al.
Pubblicazione: (2025)
di: Cheng, Myra, et al.
Pubblicazione: (2025)
Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs
di: de Landa, Joseba Fernandez, et al.
Pubblicazione: (2026)
di: de Landa, Joseba Fernandez, et al.
Pubblicazione: (2026)
Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
di: Choi, Sooyung, et al.
Pubblicazione: (2025)
Epistemic Constitutionalism Or: how to avoid coherence bias
di: Loi, Michele
Pubblicazione: (2026)
di: Loi, Michele
Pubblicazione: (2026)
The Polite Liar: Epistemic Pathology in Language Models
di: DeVilling, Bentley
Pubblicazione: (2025)
di: DeVilling, Bentley
Pubblicazione: (2025)
Speak Easy: Eliciting Harmful Jailbreaks from LLMs with Simple Interactions
di: Chan, Yik Siu, et al.
Pubblicazione: (2025)
di: Chan, Yik Siu, et al.
Pubblicazione: (2025)
The AI Cognitive Trojan Horse: How Large Language Models May Bypass Human Epistemic Vigilance
di: Maynard, Andrew D.
Pubblicazione: (2026)
di: Maynard, Andrew D.
Pubblicazione: (2026)
The Illusion of Friendship: Why Generative AI Demands Unprecedented Ethical Vigilance
di: Islam, Md Zahidul
Pubblicazione: (2026)
di: Islam, Md Zahidul
Pubblicazione: (2026)
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
di: Kabir, Mohsinul, et al.
Pubblicazione: (2025)
The Earth is Flat because...: Investigating LLMs' Belief towards Misinformation via Persuasive Conversation
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
di: Xu, Rongwu, et al.
Pubblicazione: (2023)
Why Slop Matters
di: Kommers, Cody, et al.
Pubblicazione: (2025)
di: Kommers, Cody, et al.
Pubblicazione: (2025)
FairBelief -- Assessing Harmful Beliefs in Language Models
di: Setzu, Mattia, et al.
Pubblicazione: (2024)
di: Setzu, Mattia, et al.
Pubblicazione: (2024)
"I Am the One and Only, Your Cyber BFF": Understanding the Impact of GenAI Requires Understanding the Impact of Anthropomorphic AI
di: Cheng, Myra, et al.
Pubblicazione: (2024)
di: Cheng, Myra, et al.
Pubblicazione: (2024)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
di: Hawkins, John, et al.
Pubblicazione: (2025)
di: Hawkins, John, et al.
Pubblicazione: (2025)
Why LLMs Fail at Causal Discovery and How Interventional Agents Escape
di: Roy, Amartya, et al.
Pubblicazione: (2026)
di: Roy, Amartya, et al.
Pubblicazione: (2026)
Noosemia: toward a Cognitive and Phenomenological Account of Intentionality Attribution in Human-Generative AI Interaction
di: De Santis, Enrico, et al.
Pubblicazione: (2025)
di: De Santis, Enrico, et al.
Pubblicazione: (2025)
Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
di: Chien, Jennifer, et al.
Pubblicazione: (2024)
Security Challenges in AI Agent Deployment: Insights from a Large Scale Public Competition
di: Zou, Andy, et al.
Pubblicazione: (2025)
di: Zou, Andy, et al.
Pubblicazione: (2025)
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs
di: Gao, Cheng, et al.
Pubblicazione: (2025)
di: Gao, Cheng, et al.
Pubblicazione: (2025)
Cultural Compass: A Framework for Organizing Societal Norms to Detect Violations in Human-AI Conversations
di: Cheng, Myra, et al.
Pubblicazione: (2026)
di: Cheng, Myra, et al.
Pubblicazione: (2026)
Look Before You Leap: Enhancing Attention and Vigilance Regarding Harmful Content with GuidelineLLM
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
di: Zhang, Shaoqing, et al.
Pubblicazione: (2024)
Harmful Suicide Content Detection
di: Park, Kyumin, et al.
Pubblicazione: (2024)
di: Park, Kyumin, et al.
Pubblicazione: (2024)
Empathy and the Right to Be an Exception: What LLMs Can and Cannot Do
di: Kidder, William, et al.
Pubblicazione: (2024)
di: Kidder, William, et al.
Pubblicazione: (2024)
Red Lines and Grey Zones in the Fog of War: Benchmarking Legal Risk, Moral Harm, and Regional Bias in Large Language Model Military Decision-Making
di: Drinkall, Toby
Pubblicazione: (2025)
di: Drinkall, Toby
Pubblicazione: (2025)
Othering and low status framing of immigrant cuisines in US restaurant reviews and large language models
di: Luo, Yiwei, et al.
Pubblicazione: (2023)
di: Luo, Yiwei, et al.
Pubblicazione: (2023)
RealHarm: A Collection of Real-World Language Model Application Failures
di: Jeune, Pierre Le, et al.
Pubblicazione: (2025)
di: Jeune, Pierre Le, et al.
Pubblicazione: (2025)
Laissez-Faire Harms: Algorithmic Biases in Generative Language Models
di: Shieh, Evan, et al.
Pubblicazione: (2024)
di: Shieh, Evan, et al.
Pubblicazione: (2024)
The 17% Gap: Quantifying Epistemic Decay in AI-Assisted Survey Papers
di: İlter, H. Kemal
Pubblicazione: (2026)
di: İlter, H. Kemal
Pubblicazione: (2026)
Documenti analoghi
-
HumT DumT: Measuring and controlling human-like language in LLMs
di: Cheng, Myra, et al.
Pubblicazione: (2025) -
AnthroScore: A Computational Linguistic Measure of Anthropomorphism
di: Cheng, Myra, et al.
Pubblicazione: (2024) -
ELEPHANT: Measuring and understanding social sycophancy in LLMs
di: Cheng, Myra, et al.
Pubblicazione: (2025) -
Verbalizing LLMs' assumptions to explain and control sycophancy
di: Cheng, Myra, et al.
Pubblicazione: (2026) -
Belief in the Machine: Investigating Epistemological Blind Spots of Language Models
di: Suzgun, Mirac, et al.
Pubblicazione: (2024)