Exploring Safety Alignment Evaluation of LLMs in Chinese Mental Health Dialogues via LLM-as-Judge
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Cai, Yunna, Wang, Fan, Wang, Haowei, Wang, Kun, Yang, Kailai, Ananiadou, Sophia, Li, Moyan, Fan, Mingming |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
par: Kabir, Mohsinul, et autres
Publié: (2025)
par: Kabir, Mohsinul, et autres
Publié: (2025)
Towards Interpretable Mental Health Analysis with Large Language Models
par: Yang, Kailai, et autres
Publié: (2023)
par: Yang, Kailai, et autres
Publié: (2023)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
par: Kabir, Mohsinul, et autres
Publié: (2026)
par: Kabir, Mohsinul, et autres
Publié: (2026)
Evaluating the Clinical Safety of LLMs in Response to High-Risk Mental Health Disclosures
par: Shah, Siddharth, et autres
Publié: (2025)
par: Shah, Siddharth, et autres
Publié: (2025)
MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models
par: Yang, Kailai, et autres
Publié: (2023)
par: Yang, Kailai, et autres
Publié: (2023)
LLM-as-a-Judge in Healthcare: A Scoping Analysis of Applications, Methods, and Human Alignment
par: Li, Lingyao, et autres
Publié: (2026)
par: Li, Lingyao, et autres
Publié: (2026)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
par: Arnaiz-Rodriguez, Adrian, et autres
Publié: (2025)
par: Arnaiz-Rodriguez, Adrian, et autres
Publié: (2025)
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation
par: Badawi, Abeer, et autres
Publié: (2025)
par: Badawi, Abeer, et autres
Publié: (2025)
Exploring LLMs for Predicting Tutor Strategy and Student Outcomes in Dialogues
par: Ikram, Fareya, et autres
Publié: (2025)
par: Ikram, Fareya, et autres
Publié: (2025)
MMAFFBen: A Multilingual and Multimodal Affective Analysis Benchmark for Evaluating LLMs and VLMs
par: Liu, Zhiwei, et autres
Publié: (2025)
par: Liu, Zhiwei, et autres
Publié: (2025)
LLM Use for Mental Health: Crowdsourcing Users' Sentiment-based Perspectives and Values from Social Discussions
par: Li, Lingyao, et autres
Publié: (2025)
par: Li, Lingyao, et autres
Publié: (2025)
Safety Evaluation and Enhancement of DeepSeek Models in Chinese Contexts
par: Zhang, Wenjing, et autres
Publié: (2025)
par: Zhang, Wenjing, et autres
Publié: (2025)
SPA: Achieving Consensus in LLM Alignment via Self-Priority Optimization
par: Huang, Yue, et autres
Publié: (2025)
par: Huang, Yue, et autres
Publié: (2025)
Judging the Judges: Human Validation of Multi-LLM Evaluation for High-Quality K--12 Science Instructional Materials
par: He, Peng, et autres
Publié: (2026)
par: He, Peng, et autres
Publié: (2026)
EmoLLMs: A Series of Emotional Large Language Models and Annotation Tools for Comprehensive Affective Analysis
par: Liu, Zhiwei, et autres
Publié: (2024)
par: Liu, Zhiwei, et autres
Publié: (2024)
LLM Safety Alignment is Divergence Estimation in Disguise
par: Haldar, Rajdeep, et autres
Publié: (2025)
par: Haldar, Rajdeep, et autres
Publié: (2025)
EmoAgent: Assessing and Safeguarding Human-AI Interaction for Mental Health Safety
par: Qiu, Jiahao, et autres
Publié: (2025)
par: Qiu, Jiahao, et autres
Publié: (2025)
Exploring Knowledge Tracing in Tutor-Student Dialogues using LLMs
par: Scarlatos, Alexander, et autres
Publié: (2024)
par: Scarlatos, Alexander, et autres
Publié: (2024)
ConspEmoLLM: Conspiracy Theory Detection Using an Emotion-Based Large Language Model
par: Liu, Zhiwei, et autres
Publié: (2024)
par: Liu, Zhiwei, et autres
Publié: (2024)
RAEmoLLM: Retrieval Augmented LLMs for Cross-Domain Misinformation Detection Using In-Context Learning Based on Emotional Information
par: Liu, Zhiwei, et autres
Publié: (2024)
par: Liu, Zhiwei, et autres
Publié: (2024)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
par: Yu, Zeping, et autres
Publié: (2025)
par: Yu, Zeping, et autres
Publié: (2025)
Rumor Detection by Multi-task Suffix Learning based on Time-series Dual Sentiments
par: Liu, Zhiwei, et autres
Publié: (2025)
par: Liu, Zhiwei, et autres
Publié: (2025)
Addressing the sustainable AI trilemma: a case study on LLM agents and RAG
par: Wu, Hui, et autres
Publié: (2025)
par: Wu, Hui, et autres
Publié: (2025)
An Offline Mobile Conversational Agent for Mental Health Support: Learning from Emotional Dialogues and Psychological Texts with Student-Centered Evaluation
par: A, Vimaleswar, et autres
Publié: (2025)
par: A, Vimaleswar, et autres
Publié: (2025)
Engagement-Optimized Care: When LLMs become Mental Health Infrastructure
par: Vecchione, Briana, et autres
Publié: (2026)
par: Vecchione, Briana, et autres
Publié: (2026)
Dean of LLM Tutors: Exploring Comprehensive and Automated Evaluation of LLM-generated Educational Feedback via LLM Feedback Evaluators
par: Qian, Keyang, et autres
Publié: (2025)
par: Qian, Keyang, et autres
Publié: (2025)
MetaAligner: Towards Generalizable Multi-Objective Alignment of Language Models
par: Yang, Kailai, et autres
Publié: (2024)
par: Yang, Kailai, et autres
Publié: (2024)
Can LLM be a Personalized Judge?
par: Dong, Yijiang River, et autres
Publié: (2024)
par: Dong, Yijiang River, et autres
Publié: (2024)
Randomness, Not Representation: The Unreliability of Evaluating Cultural Alignment in LLMs
par: Khan, Ariba, et autres
Publié: (2025)
par: Khan, Ariba, et autres
Publié: (2025)
Responsible Evaluation of AI for Mental Health
par: Arnaout, Hiba, et autres
Publié: (2026)
par: Arnaout, Hiba, et autres
Publié: (2026)
SouLLMate: An Adaptive LLM-Driven System for Advanced Mental Health Support and Assessment, Based on a Systematic Application Survey
par: Guo, Qiming, et autres
Publié: (2024)
par: Guo, Qiming, et autres
Publié: (2024)
Beyond Simulations: What 20,000 Real Conversations Reveal About Mental Health AI Safety
par: Stamatis, Caitlin A., et autres
Publié: (2026)
par: Stamatis, Caitlin A., et autres
Publié: (2026)
The Jade Gateway to Trust: Exploring How Socio-Cultural Perspectives Shape Trust Within Chinese NFT Communities
par: Cao, Yi-Fan, et autres
Publié: (2025)
par: Cao, Yi-Fan, et autres
Publié: (2025)
The Geography of Transportation Cybersecurity: Visitor Flows, Industry Clusters, and Spatial Dynamics
par: Wang, Yuhao, et autres
Publié: (2025)
par: Wang, Yuhao, et autres
Publié: (2025)
"It Listens Better Than My Therapist": Exploring Social Media Discourse on LLMs as Mental Health Tool
par: Haensch, Anna-Carolina
Publié: (2025)
par: Haensch, Anna-Carolina
Publié: (2025)
Exploring Socio-Cultural Challenges and Opportunities in Designing Mental Health Chatbots for Adolescents in India
par: Sehgal, Neil K. R., et autres
Publié: (2025)
par: Sehgal, Neil K. R., et autres
Publié: (2025)
What's on Your Mind? Exploring Privacy of Mental Health Apps
par: Georgiou, Chloe, et autres
Publié: (2026)
par: Georgiou, Chloe, et autres
Publié: (2026)
MoPHES:Leveraging on-device LLMs as Agent for Mobile Psychological Health Evaluation and Support
par: Wei, Xun, et autres
Publié: (2025)
par: Wei, Xun, et autres
Publié: (2025)
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
par: Yu, Zeping, et autres
Publié: (2024)
par: Yu, Zeping, et autres
Publié: (2024)
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
par: Kim, Jiseon, et autres
Publié: (2025)
par: Kim, Jiseon, et autres
Publié: (2025)
Documents similaires
-
Break the Checkbox: Challenging Closed-Style Evaluations of Cultural Alignment in LLMs
par: Kabir, Mohsinul, et autres
Publié: (2025) -
Towards Interpretable Mental Health Analysis with Large Language Models
par: Yang, Kailai, et autres
Publié: (2023) -
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
par: Kabir, Mohsinul, et autres
Publié: (2026) -
Evaluating the Clinical Safety of LLMs in Response to High-Risk Mental Health Disclosures
par: Shah, Siddharth, et autres
Publié: (2025) -
MentaLLaMA: Interpretable Mental Health Analysis on Social Media with Large Language Models
par: Yang, Kailai, et autres
Publié: (2023)