Auditing Stealth Sycophancy in Mental-Health Dialogue: Structured Clinical-State Diagnostics and Clean Matched Benchmarks
Fuente:
arXiv
Saved in:
| Main Authors: | Han, Tianze, Xu, Beining, Zhang, Hanbo, Lu, Yongming |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TECP: Token-Entropy Conformal Prediction for LLMs
by: Xu, Beining, et al.
Published: (2025)
by: Xu, Beining, et al.
Published: (2025)
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026)
by: Tu, Xinming, et al.
Published: (2026)
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
by: Natan, Shahar Ben, et al.
Published: (2026)
by: Natan, Shahar Ben, et al.
Published: (2026)
OnRL-RAG: Real-Time Personalized Mental Health Dialogue System
by: Bilal, Ahsan, et al.
Published: (2025)
by: Bilal, Ahsan, et al.
Published: (2025)
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
by: Zhou, Wenrui, et al.
Published: (2025)
by: Zhou, Wenrui, et al.
Published: (2025)
An Ubuntu-Guided Large Language Model Framework for Cognitive Behavioral Mental Health Dialogue
by: Forane, Sontaga G., et al.
Published: (2026)
by: Forane, Sontaga G., et al.
Published: (2026)
The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
by: Cheng, Youyou, et al.
Published: (2026)
by: Cheng, Youyou, et al.
Published: (2026)
BrokenMath: A Benchmark for Sycophancy in Theorem Proving with LLMs
by: Petrov, Ivo, et al.
Published: (2025)
by: Petrov, Ivo, et al.
Published: (2025)
BASIL: Bayesian Assessment of Sycophancy in LLMs
by: Atwell, Katherine, et al.
Published: (2025)
by: Atwell, Katherine, et al.
Published: (2025)
Sycophancy Hides Linearly in the Attention Heads
by: Genadi, Rifo, et al.
Published: (2026)
by: Genadi, Rifo, et al.
Published: (2026)
Adapting Text-based Dialogue State Tracker for Spoken Dialogues
by: Yoon, Jaeseok, et al.
Published: (2023)
by: Yoon, Jaeseok, et al.
Published: (2023)
From Medical Records to Diagnostic Dialogues: A Clinical-Grounded Approach and Dataset for Psychiatric Comorbidity
by: Wan, Tianxi, et al.
Published: (2025)
by: Wan, Tianxi, et al.
Published: (2025)
Sycophancy in Large Language Models: Causes and Mitigations
by: Malmqvist, Lars
Published: (2024)
by: Malmqvist, Lars
Published: (2024)
Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment
by: Yeh, Samuel, et al.
Published: (2025)
by: Yeh, Samuel, et al.
Published: (2025)
A Comprehensive Review of Datasets for Clinical Mental Health AI Systems
by: Mandal, Aishik, et al.
Published: (2025)
by: Mandal, Aishik, et al.
Published: (2025)
Acting Flatterers via LLMs Sycophancy: Combating Clickbait with LLMs Opposing-Stance Reasoning
by: Zhang, Chaowei, et al.
Published: (2026)
by: Zhang, Chaowei, et al.
Published: (2026)
Do Physicians Know How to Prompt? The Need for Automatic Prompt Optimization Help in Clinical Note Generation
by: Yao, Zonghai, et al.
Published: (2023)
by: Yao, Zonghai, et al.
Published: (2023)
MindEval: Benchmarking Language Models on Multi-turn Mental Health Support
by: Pombal, José, et al.
Published: (2025)
by: Pombal, José, et al.
Published: (2025)
Reasoning Isn't Enough: Examining Truth-Bias and Sycophancy in LLMs
by: Barkett, Emilio, et al.
Published: (2025)
by: Barkett, Emilio, et al.
Published: (2025)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024)
by: Denison, Carson, et al.
Published: (2024)
TrustMH-Bench: A Comprehensive Benchmark for Evaluating the Trustworthiness of Large Language Models in Mental Health
by: Xiong, Zixin, et al.
Published: (2026)
by: Xiong, Zixin, et al.
Published: (2026)
Connecting the Dots: Benchmarking Reflective Memory in Long-Horizon Dialogue
by: Lin, Jingjie, et al.
Published: (2026)
by: Lin, Jingjie, et al.
Published: (2026)
AgentStealth: Reinforcing Large Language Model for Anonymizing User-generated Text
by: Shao, Chenyang, et al.
Published: (2025)
by: Shao, Chenyang, et al.
Published: (2025)
Sycophancy in Vision-Language Models: A Systematic Analysis and an Inference-Time Mitigation Framework
by: Zhao, Yunpu, et al.
Published: (2024)
by: Zhao, Yunpu, et al.
Published: (2024)
From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making
by: Jain, Raunak
Published: (2026)
by: Jain, Raunak
Published: (2026)
Reasoning Like a Doctor: Improving Medical Dialogue Systems via Diagnostic Reasoning Process Alignment
by: Xu, Kaishuai, et al.
Published: (2024)
by: Xu, Kaishuai, et al.
Published: (2024)
ContextGuard: Structured Self-Auditing for Context Learning in Language Models
by: Jin, Hongbo, et al.
Published: (2026)
by: Jin, Hongbo, et al.
Published: (2026)
MIRA: A Bilingual Benchmark for Medical Information Response Audit
by: Xu, Mengyu, et al.
Published: (2026)
by: Xu, Mengyu, et al.
Published: (2026)
Accounting for Sycophancy in Language Model Uncertainty Estimation
by: Sicilia, Anthony, et al.
Published: (2024)
by: Sicilia, Anthony, et al.
Published: (2024)
EthicsMH: A Pilot Benchmark for Ethical Reasoning in Mental Health AI
by: Kasu, Sai Kartheek Reddy
Published: (2025)
by: Kasu, Sai Kartheek Reddy
Published: (2025)
MindSET: Advancing Mental Health Benchmarking through Large-Scale Social Media Data
by: Mankarious, Saad, et al.
Published: (2025)
by: Mankarious, Saad, et al.
Published: (2025)
Intent-driven In-context Learning for Few-shot Dialogue State Tracking
by: Yi, Zihao, et al.
Published: (2024)
by: Yi, Zihao, et al.
Published: (2024)
Beacon: Single-Turn Diagnosis and Mitigation of Latent Sycophancy in Large Language Models
by: Pandey, Sanskar, et al.
Published: (2025)
by: Pandey, Sanskar, et al.
Published: (2025)
A Dual-Prompting for Interpretable Mental Health Language Models
by: Jeon, Hyolim, et al.
Published: (2024)
by: Jeon, Hyolim, et al.
Published: (2024)
ToMATO: Verbalizing the Mental States of Role-Playing LLMs for Benchmarking Theory of Mind
by: Shinoda, Kazutoshi, et al.
Published: (2025)
by: Shinoda, Kazutoshi, et al.
Published: (2025)
Exploration of Perceptual Speech Features for Clinical Decision-Support in Mental Health Care
by: Lyberatos, Vassilis, et al.
Published: (2026)
by: Lyberatos, Vassilis, et al.
Published: (2026)
Improving Phishing Email Detection Performance of Small Large Language Models
by: Lin, Zijie, et al.
Published: (2025)
by: Lin, Zijie, et al.
Published: (2025)
ClinicalLab: Aligning Agents for Multi-Departmental Clinical Diagnostics in the Real World
by: Yan, Weixiang, et al.
Published: (2024)
by: Yan, Weixiang, et al.
Published: (2024)
MONICA: Real-Time Monitoring and Calibration of Chain-of-Thought Sycophancy in Large Reasoning Models
by: Hu, Jingyu, et al.
Published: (2025)
by: Hu, Jingyu, et al.
Published: (2025)
SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization
by: Mishra, Prakamya, et al.
Published: (2024)
by: Mishra, Prakamya, et al.
Published: (2024)
Similar Items
-
TECP: Token-Entropy Conformal Prediction for LLMs
by: Xu, Beining, et al.
Published: (2025) -
BenchGuard: Who Guards the Benchmarks? Automated Auditing of LLM Agent Benchmarks
by: Tu, Xinming, et al.
Published: (2026) -
Not Your Typical Sycophant: The Elusive Nature of Sycophancy in Large Language Models
by: Natan, Shahar Ben, et al.
Published: (2026) -
OnRL-RAG: Real-Time Personalized Mental Health Dialogue System
by: Bilal, Ahsan, et al.
Published: (2025) -
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs
by: Zhou, Wenrui, et al.
Published: (2025)