LLM Spirals of Delusion: A Benchmarking Audit Study of AI Chatbot Interfaces
Fuente:
arXiv
Saved in:
| Main Authors: | Kirgis, Peter, Hawriluk, Ben, Feng, Sherrie, Bilimer, Aslan, Paech, Sam, Tufekci, Zeynep |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Dynamics of Delusion: Modeling Bidirectional False Belief Amplification in Human-Chatbot Dialogue
by: Mehta, Ashish, et al.
Published: (2026)
by: Mehta, Ashish, et al.
Published: (2026)
From Chatbots to Confidants: A Cross-Cultural Study of LLM Adoption for Emotional Support
by: Amat-Lefort, Natalia, et al.
Published: (2026)
by: Amat-Lefort, Natalia, et al.
Published: (2026)
Can AI Have a Personality? Prompt Engineering for AI Personality Simulation: A Chatbot Case Study in Gender-Affirming Voice Therapy Training
by: Jackson, Tailon D., et al.
Published: (2025)
by: Jackson, Tailon D., et al.
Published: (2025)
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
by: Shaikh, Omar, et al.
Published: (2025)
by: Shaikh, Omar, et al.
Published: (2025)
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)
by: Yoon, Sion, et al.
Published: (2024)
AI Psychosis: Does Conversational AI Amplify Delusion-Related Language?
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
by: Shimgekar, Soorya Ram, et al.
Published: (2026)
Online vs Offline: A Comparative Study of First-Party and Third-Party Evaluations of Social Chatbots
by: Svikhnushina, Ekaterina, et al.
Published: (2024)
by: Svikhnushina, Ekaterina, et al.
Published: (2024)
Understanding Learner-LLM Chatbot Interactions and the Impact of Prompting Guidelines
by: Koyuturk, Cansu, et al.
Published: (2025)
by: Koyuturk, Cansu, et al.
Published: (2025)
Dialogue Act Patterns in GenAI-Mediated L2 Oral Practice: A Sequential Analysis of Learner-Chatbot Interactions
by: He, Liqun, et al.
Published: (2026)
by: He, Liqun, et al.
Published: (2026)
The Impact of a Chatbot's Ephemerality-Framing on Self-Disclosure Perceptions
by: Cox, Samuel Rhys, et al.
Published: (2025)
by: Cox, Samuel Rhys, et al.
Published: (2025)
Benchmarking LLM Tool-Use in the Wild
by: Yu, Peijie, et al.
Published: (2026)
by: Yu, Peijie, et al.
Published: (2026)
Low-code LLM: Graphical User Interface over Large Language Models
by: Cai, Yuzhe, et al.
Published: (2023)
by: Cai, Yuzhe, et al.
Published: (2023)
Polite But Boring? Trade-offs Between Engagement and Psychological Reactance to Chatbot Feedback Styles
by: Cox, Samuel Rhys, et al.
Published: (2026)
by: Cox, Samuel Rhys, et al.
Published: (2026)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
PaperTrail: A Claim-Evidence Interface for Grounding Provenance in LLM-based Scholarly Q&A
by: Martin-Boyle, Anna, et al.
Published: (2026)
by: Martin-Boyle, Anna, et al.
Published: (2026)
Towards an LLM-Based Speech Interface for Robot-Assisted Feeding
by: Yuan, Jessie, et al.
Published: (2024)
by: Yuan, Jessie, et al.
Published: (2024)
A Piece of Theatre: Investigating How Teachers Design LLM Chatbots to Assist Adolescent Cyberbullying Education
by: Hedderich, Michael A., et al.
Published: (2024)
by: Hedderich, Michael A., et al.
Published: (2024)
Human-Centred LLM Privacy Audits: Findings and Frictions
by: Staufer, Dimitri, et al.
Published: (2026)
by: Staufer, Dimitri, et al.
Published: (2026)
Supporting Student Decisions on Learning Recommendations: An LLM-Based Chatbot with Knowledge Graph Contextualization for Conversational Explainability and Mentoring
by: Abu-Rasheed, Hasan, et al.
Published: (2024)
by: Abu-Rasheed, Hasan, et al.
Published: (2024)
Comparing How a Chatbot References User Utterances from Previous Chatting Sessions: An Investigation of Users' Privacy Concerns and Perceptions
by: Cox, Samuel Rhys, et al.
Published: (2023)
by: Cox, Samuel Rhys, et al.
Published: (2023)
Prototypical Human-AI Collaboration Behaviors from LLM-Assisted Writing in the Wild
by: Mysore, Sheshera, et al.
Published: (2025)
by: Mysore, Sheshera, et al.
Published: (2025)
Many Ways to Be Fake: Benchmarking Fake News Detection Under Strategy-Driven AI Generation
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Vibe Coding, Interface Flattening
by: Jin, Hongrui
Published: (2025)
by: Jin, Hongrui
Published: (2025)
From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered Analysis
by: Li, Zhuoyan, et al.
Published: (2025)
by: Li, Zhuoyan, et al.
Published: (2025)
Cross-Lingual Sentiment Misalignment: Auditing Multilingual Language Models for Inversion Risk, Dialectal Representation, and Affective Stability
by: Lia, Nusrat Jahan, et al.
Published: (2026)
by: Lia, Nusrat Jahan, et al.
Published: (2026)
MIRAGE: Multi-model Interface for Reviewing and Auditing Generative Text-to-Image AI
by: Maldaner, Matheus Kunzler, et al.
Published: (2025)
by: Maldaner, Matheus Kunzler, et al.
Published: (2025)
Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
by: Chandra, Kartik, et al.
Published: (2026)
by: Chandra, Kartik, et al.
Published: (2026)
ShareChat: A Dataset of Chatbot Conversations in the Wild
by: Yan, Yueru, et al.
Published: (2025)
by: Yan, Yueru, et al.
Published: (2025)
Telephone Surveys Meet Conversational AI: Evaluating a LLM-Based Telephone Survey System at Scale
by: Lang, Max M., et al.
Published: (2025)
by: Lang, Max M., et al.
Published: (2025)
Sniff AI: Is My 'Spicy' Your 'Spicy'? Exploring LLM's Perceptual Alignment with Human Smell Experiences
by: Zhong, Shu, et al.
Published: (2024)
by: Zhong, Shu, et al.
Published: (2024)
AI vs. Human Judgment of Content Moderation: LLM-as-a-Judge and Ethics-Based Response Refusals
by: Pasch, Stefan
Published: (2025)
by: Pasch, Stefan
Published: (2025)
Evaluating LLM-Generated Q&A Test: a Student-Centered Study
by: Wróblewska, Anna, et al.
Published: (2025)
by: Wróblewska, Anna, et al.
Published: (2025)
A Comparative Study on Annotation Quality of Crowdsourcing and LLM via Label Aggregation
by: Li, Jiyi
Published: (2024)
by: Li, Jiyi
Published: (2024)
The Persuasion Paradox: When LLM Explanations Fail to Improve Human-AI Team Performance
by: Cohen, Ruth, et al.
Published: (2026)
by: Cohen, Ruth, et al.
Published: (2026)
Exploring the Efficacy of Large Language Models in Summarizing Mental Health Counseling Sessions: A Benchmark Study
by: Adhikary, Prottay Kumar, et al.
Published: (2024)
by: Adhikary, Prottay Kumar, et al.
Published: (2024)
A Framework for Evaluating Appropriateness, Trustworthiness, and Safety in Mental Wellness AI Chatbots
by: Chen, Lucia, et al.
Published: (2024)
by: Chen, Lucia, et al.
Published: (2024)
ConsistencyAI: A Benchmark to Assess LLMs' Factual Consistency When Responding to Different Demographic Groups
by: Banyas, Peter, et al.
Published: (2025)
by: Banyas, Peter, et al.
Published: (2025)
Media of Langue: The Interface for Exploring Word Translation Network/Space
by: Muramoto, Goki, et al.
Published: (2023)
by: Muramoto, Goki, et al.
Published: (2023)
What Makes LLM Agent Simulations Useful for Policy Practice? An Iterative Design Study in Emergency Preparedness
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Building Trust in Mental Health Chatbots: Safety Metrics and LLM-Based Evaluation Tools
by: Park, Jung In, et al.
Published: (2024)
by: Park, Jung In, et al.
Published: (2024)
Similar Items
-
The Dynamics of Delusion: Modeling Bidirectional False Belief Amplification in Human-Chatbot Dialogue
by: Mehta, Ashish, et al.
Published: (2026) -
From Chatbots to Confidants: A Cross-Cultural Study of LLM Adoption for Emotional Support
by: Amat-Lefort, Natalia, et al.
Published: (2026) -
Can AI Have a Personality? Prompt Engineering for AI Personality Simulation: A Chatbot Case Study in Gender-Affirming Voice Therapy Training
by: Jackson, Tailon D., et al.
Published: (2025) -
Navigating Rifts in Human-LLM Grounding: Study and Benchmark
by: Shaikh, Omar, et al.
Published: (2025) -
Designing and Evaluating Multi-Chatbot Interface for Human-AI Communication: Preliminary Findings from a Persuasion Task
by: Yoon, Sion, et al.
Published: (2024)