Unintended Harms of Value-Aligned LLMs: Psychological and Empirical Insights
Fuente:
arXiv
Saved in:
| Main Authors: | Choi, Sooyung, Lee, Jaehyeok, Yi, Xiaoyuan, Yao, Jing, Xie, Xing, Bak, JinYeong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026)
by: Lee, Jaehyeok, et al.
Published: (2026)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
by: Lee, Jaehyeok, et al.
Published: (2024)
by: Lee, Jaehyeok, et al.
Published: (2024)
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
by: Kim, HyunJin, et al.
Published: (2025)
by: Kim, HyunJin, et al.
Published: (2025)
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
by: Yao, Jing, et al.
Published: (2025)
by: Yao, Jing, et al.
Published: (2025)
Harmful Suicide Content Detection
by: Park, Kyumin, et al.
Published: (2024)
by: Park, Kyumin, et al.
Published: (2024)
MentalAgora: A Gateway to Advanced Personalized Care in Mental Health through Multi-Agent Debating and Attribute Control
by: Lee, Yeonji, et al.
Published: (2024)
by: Lee, Yeonji, et al.
Published: (2024)
MoVa: Towards Generalizable Classification of Human Morals and Values
by: Chen, Ziyu, et al.
Published: (2025)
by: Chen, Ziyu, et al.
Published: (2025)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
KpopMT: Translation Dataset with Terminology for Kpop Fandom
by: Kim, JiWoo, et al.
Published: (2024)
by: Kim, JiWoo, et al.
Published: (2024)
Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions
by: Kim, JiWoo, et al.
Published: (2025)
by: Kim, JiWoo, et al.
Published: (2025)
PEMA: An Offsite-Tunable Plug-in External Memory Adaptation for Language Models
by: Kim, HyunJin, et al.
Published: (2023)
by: Kim, HyunJin, et al.
Published: (2023)
The Incomplete Bridge: How AI Research (Mis)Engages with Psychology
by: Jiang, Han, et al.
Published: (2025)
by: Jiang, Han, et al.
Published: (2025)
Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2026)
by: Tan, Bryan Chen Zhengyu, et al.
Published: (2026)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
by: Jiang, Han, et al.
Published: (2024)
by: Jiang, Han, et al.
Published: (2024)
Denevil: Towards Deciphering and Navigating the Ethical Values of Large Language Models via Instruction Learning
by: Duan, Shitong, et al.
Published: (2023)
by: Duan, Shitong, et al.
Published: (2023)
CLAVE: An Adaptive Framework for Evaluating Values of LLM Generated Responses
by: Yao, Jing, et al.
Published: (2024)
by: Yao, Jing, et al.
Published: (2024)
Expected Harm: Rethinking Safety Evaluation of (Mis)Aligned LLMs
by: Chen, Yen-Shan, et al.
Published: (2026)
by: Chen, Yen-Shan, et al.
Published: (2026)
Memoria: Resolving Fateful Forgetting Problem through Human-Inspired Memory Architecture
by: Park, Sangjun, et al.
Published: (2023)
by: Park, Sangjun, et al.
Published: (2023)
Diverse, but Divisive: LLMs Can Exaggerate Gender Differences in Opinion Related to Harms of Misinformation
by: Neumann, Terrence, et al.
Published: (2024)
by: Neumann, Terrence, et al.
Published: (2024)
On the Sensitivity of Instruction-tuned LLMs to Harmful Sentences in Long Inputs
by: Ghorbanpour, Faeze, et al.
Published: (2025)
by: Ghorbanpour, Faeze, et al.
Published: (2025)
CDEval: A Benchmark for Measuring the Cultural Dimensions of Large Language Models
by: Wang, Yuhang, et al.
Published: (2023)
by: Wang, Yuhang, et al.
Published: (2023)
Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
by: Wang, Xingqi, et al.
Published: (2024)
by: Wang, Xingqi, et al.
Published: (2024)
Between Help and Harm: An Evaluation of Mental Health Crisis Handling by LLMs
by: Arnaiz-Rodriguez, Adrian, et al.
Published: (2025)
by: Arnaiz-Rodriguez, Adrian, et al.
Published: (2025)
Do Prevalent Bias Metrics Capture Allocational Harms from LLMs?
by: Cyberey, Hannah, et al.
Published: (2024)
by: Cyberey, Hannah, et al.
Published: (2024)
Careless Whisper: Speech-to-Text Hallucination Harms
by: Koenecke, Allison, et al.
Published: (2024)
by: Koenecke, Allison, et al.
Published: (2024)
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
by: Li, Jing-Jing, et al.
Published: (2026)
by: Li, Jing-Jing, et al.
Published: (2026)
IROTE: Human-like Traits Elicitation of Large Language Model via In-Context Self-Reflective Optimization
by: Bai, Yuzhuo, et al.
Published: (2025)
by: Bai, Yuzhuo, et al.
Published: (2025)
Unintended Impacts of LLM Alignment on Global Representation
by: Ryan, Michael J., et al.
Published: (2024)
by: Ryan, Michael J., et al.
Published: (2024)
Translating Hanja Historical Documents to Contemporary Korean and English
by: Son, Juhee, et al.
Published: (2022)
by: Son, Juhee, et al.
Published: (2022)
HRIPBench: Benchmarking LLMs in Harm Reduction Information Provision to Support People Who Use Drugs
by: Wang, Kaixuan, et al.
Published: (2025)
by: Wang, Kaixuan, et al.
Published: (2025)
Leveraging Implicit Sentiments: Enhancing Reliability and Validity in Psychological Trait Evaluation of LLMs
by: Ma, Huanhuan, et al.
Published: (2025)
by: Ma, Huanhuan, et al.
Published: (2025)
Harmful Speech Detection by Language Models Exhibits Gender-Queer Dialect Bias
by: Dorn, Rebecca, et al.
Published: (2024)
by: Dorn, Rebecca, et al.
Published: (2024)
Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM
by: Zhang, Chi, et al.
Published: (2025)
by: Zhang, Chi, et al.
Published: (2025)
Survival at Any Cost? LLMs and the Choice Between Self-Preservation and Human Harm
by: Mohamadi, Alireza, et al.
Published: (2025)
by: Mohamadi, Alireza, et al.
Published: (2025)
CASHG: Context-Aware Stylized Online Handwriting Generation
by: Shin, Jinsu, et al.
Published: (2026)
by: Shin, Jinsu, et al.
Published: (2026)
Are Social Sentiments Inherent in LLMs? An Empirical Study on Extraction of Inter-demographic Sentiments
by: Tanaka, Kunitomo, et al.
Published: (2024)
by: Tanaka, Kunitomo, et al.
Published: (2024)
Cultural Value Differences of LLMs: Prompt, Language, and Model Size
by: Zhong, Qishuai, et al.
Published: (2024)
by: Zhong, Qishuai, et al.
Published: (2024)
Somatic in the East, Psychological in the West?: Investigating Clinically-Grounded Cross-Cultural Depression Symptom Expression in LLMs
by: Sakai, Shintaro, et al.
Published: (2025)
by: Sakai, Shintaro, et al.
Published: (2025)
Accommodation and Epistemic Vigilance: A Pragmatic Account of Why LLMs Fail to Challenge Harmful Beliefs
by: Cheng, Myra, et al.
Published: (2026)
by: Cheng, Myra, et al.
Published: (2026)
From Data to Behavior: Predicting Unintended Model Behaviors Before Training
by: Wang, Mengru, et al.
Published: (2026)
by: Wang, Mengru, et al.
Published: (2026)
Similar Items
-
Distributional Open-Ended Evaluation of LLM Cultural Value Alignment Based on Value Codebook
by: Lee, Jaehyeok, et al.
Published: (2026) -
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
by: Lee, Jaehyeok, et al.
Published: (2024) -
Research Superalignment Should Advance Now with Alternating Competence and Conformity Optimization
by: Kim, HyunJin, et al.
Published: (2025) -
AdAEM: An Adaptively and Automated Extensible Measurement of LLMs' Value Difference
by: Yao, Jing, et al.
Published: (2025) -
Harmful Suicide Content Detection
by: Park, Kyumin, et al.
Published: (2024)