Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Kwon, Jea, Vecchietti, Luiz Felipe, Park, Sungwon, Cha, Meeyoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
by: Kim, Jiseon, et al.
Published: (2025)
by: Kim, Jiseon, et al.
Published: (2025)
Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions
by: Kim, Jiseon, et al.
Published: (2026)
by: Kim, Jiseon, et al.
Published: (2026)
Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
by: Park, Sungwon, et al.
Published: (2024)
by: Park, Sungwon, et al.
Published: (2024)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
by: Lee, Donggyu, et al.
Published: (2025)
by: Lee, Donggyu, et al.
Published: (2025)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
by: Kim, Minsung, et al.
Published: (2025)
by: Kim, Minsung, et al.
Published: (2025)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
by: Hu, Yueqing, et al.
Published: (2026)
by: Hu, Yueqing, et al.
Published: (2026)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
by: Vida, Karina, et al.
Published: (2024)
by: Vida, Karina, et al.
Published: (2024)
Societal Alignment Frameworks Can Improve LLM Alignment
by: Stańczak, Karolina, et al.
Published: (2025)
by: Stańczak, Karolina, et al.
Published: (2025)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
by: Backmann, Steffen, et al.
Published: (2025)
by: Backmann, Steffen, et al.
Published: (2025)
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
by: Yao, Xintong
Published: (2026)
by: Yao, Xintong
Published: (2026)
Does Cross-Cultural Alignment Change the Commonsense Morality of Language Models?
by: Jinnai, Yuu
Published: (2024)
by: Jinnai, Yuu
Published: (2024)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
by: Qu, Jinxian, et al.
Published: (2026)
by: Qu, Jinxian, et al.
Published: (2026)
Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
by: Koorndijk, Jeanice
Published: (2025)
by: Koorndijk, Jeanice
Published: (2025)
Moral Mazes in the Era of LLMs
by: Nguyen, Dang, et al.
Published: (2026)
by: Nguyen, Dang, et al.
Published: (2026)
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
by: Sauter, Adrian, et al.
Published: (2026)
by: Sauter, Adrian, et al.
Published: (2026)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Poverty mapping in Mongolia with AI-based Ger detection reveals urban slums persist after the COVID-19 pandemic
by: Yang, Jeasurk, et al.
Published: (2024)
by: Yang, Jeasurk, et al.
Published: (2024)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
by: Schoenegger, Philipp, et al.
Published: (2024)
by: Schoenegger, Philipp, et al.
Published: (2024)
ProgressGym: Alignment with a Millennium of Moral Progress
by: Qiu, Tianyi, et al.
Published: (2024)
by: Qiu, Tianyi, et al.
Published: (2024)
Scopes of Alignment
by: Varshney, Kush R., et al.
Published: (2025)
by: Varshney, Kush R., et al.
Published: (2025)
Evaluating LLM Behavior in Hiring: Implicit Weights, Fairness Across Groups, and Alignment with Human Preferences
by: Hoffmann, Morgane, et al.
Published: (2026)
by: Hoffmann, Morgane, et al.
Published: (2026)
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
by: Biancotti, Claudia, et al.
Published: (2024)
by: Biancotti, Claudia, et al.
Published: (2024)
Moral Alignment for LLM Agents
by: Tennant, Elizaveta, et al.
Published: (2024)
by: Tennant, Elizaveta, et al.
Published: (2024)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
by: Fernandes, Gustavo Lúcius, et al.
Published: (2026)
Culturally Adaptive Explainable LLM Assessment for Multilingual Information Disorder: A Human-in-the-Loop Approach
by: Jouneghani, Maziar Kianimoghadam
Published: (2026)
by: Jouneghani, Maziar Kianimoghadam
Published: (2026)
Human Preferences for Constructive Interactions in Language Model Alignment
by: Kyrychenko, Yara, et al.
Published: (2025)
by: Kyrychenko, Yara, et al.
Published: (2025)
Attributions toward Artificial Agents in a modified Moral Turing Test
by: Aharoni, Eyal, et al.
Published: (2024)
by: Aharoni, Eyal, et al.
Published: (2024)
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
by: Park, Seongchan, et al.
Published: (2026)
by: Park, Seongchan, et al.
Published: (2026)
Generalizable Slum Detection from Satellite Imagery with Mixture-of-Experts
by: Lee, Sumin, et al.
Published: (2025)
by: Lee, Sumin, et al.
Published: (2025)
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
by: Lin, Xiao, et al.
Published: (2025)
by: Lin, Xiao, et al.
Published: (2025)
A Cross-Cultural Assessment of Human Ability to Detect LLM-Generated Fake News about South Africa
by: Schlippe, Tim, et al.
Published: (2025)
by: Schlippe, Tim, et al.
Published: (2025)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
by: Costa, Davi Bastos, et al.
Published: (2025)
by: Costa, Davi Bastos, et al.
Published: (2025)
From Black-Box Confidence to Measurable Trust in Clinical AI: A Framework for Evidence, Supervision, and Staged Autonomy
by: Zabolotnii, Serhii, et al.
Published: (2026)
by: Zabolotnii, Serhii, et al.
Published: (2026)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
by: Ding, Junchen, et al.
Published: (2025)
by: Ding, Junchen, et al.
Published: (2025)
GeoSEE: Regional Socio-Economic Estimation With a Large Language Model
by: Han, Sungwon, et al.
Published: (2024)
by: Han, Sungwon, et al.
Published: (2024)
Fluent but Foreign: Even Regional LLMs Lack Cultural Alignment
by: Agarwal, Dhruv, et al.
Published: (2025)
by: Agarwal, Dhruv, et al.
Published: (2025)
In Silico Sociology: Forecasting COVID-19 Polarization with Large Language Models
by: Kozlowski, Austin C., et al.
Published: (2024)
by: Kozlowski, Austin C., et al.
Published: (2024)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
by: Bojic, Ljubisa, et al.
Published: (2026)
by: Bojic, Ljubisa, et al.
Published: (2026)
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
by: Zhou, Zhenhong, et al.
Published: (2024)
by: Zhou, Zhenhong, et al.
Published: (2024)
Similar Items
-
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
by: Kim, Jiseon, et al.
Published: (2025) -
Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions
by: Kim, Jiseon, et al.
Published: (2026) -
Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
by: Park, Sungwon, et al.
Published: (2024) -
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
by: Lee, Donggyu, et al.
Published: (2025) -
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
by: Kim, Minsung, et al.
Published: (2025)