Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Jiseon, Kwon, Jea, Vecchietti, Luiz Felipe, Dong, Wenchao, Kim, Jaehong, Cha, Meeyoung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
por: Kim, Jiseon, et al.
Publicado: (2025)
por: Kim, Jiseon, et al.
Publicado: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
por: Kwon, Jea, et al.
Publicado: (2025)
por: Kwon, Jea, et al.
Publicado: (2025)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
por: Kim, Minsung, et al.
Publicado: (2025)
por: Kim, Minsung, et al.
Publicado: (2025)
I Am Not Them: Fluid Identities and Persistent Out-group Bias in Large Language Models
por: Dong, Wenchao, et al.
Publicado: (2024)
por: Dong, Wenchao, et al.
Publicado: (2024)
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
por: Park, Seongchan, et al.
Publicado: (2026)
por: Park, Seongchan, et al.
Publicado: (2026)
Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models
por: Fränken, Jan-Philipp, et al.
Publicado: (2024)
por: Fränken, Jan-Philipp, et al.
Publicado: (2024)
Persona Setting Pitfall: Persistent Outgroup Biases in Large Language Models Arising from Social Identity Adoption
por: Dong, Wenchao, et al.
Publicado: (2024)
por: Dong, Wenchao, et al.
Publicado: (2024)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
por: Oh, Juhyun, et al.
Publicado: (2024)
por: Oh, Juhyun, et al.
Publicado: (2024)
Bilinear representation mitigates reversal curse and enables consistent model editing
por: Kim, Dong-Kyum, et al.
Publicado: (2025)
por: Kim, Dong-Kyum, et al.
Publicado: (2025)
The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas
por: Marraffini, Giovanni Franco Gabriel, et al.
Publicado: (2025)
por: Marraffini, Giovanni Franco Gabriel, et al.
Publicado: (2025)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
por: Lee, Donggyu, et al.
Publicado: (2025)
por: Lee, Donggyu, et al.
Publicado: (2025)
Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
por: Yang, Nakyeong, et al.
Publicado: (2025)
por: Yang, Nakyeong, et al.
Publicado: (2025)
How You Ask Matters! Adaptive RAG Robustness to Query Variations
por: Jang, Yunah, et al.
Publicado: (2026)
por: Jang, Yunah, et al.
Publicado: (2026)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
por: Ding, Junchen, et al.
Publicado: (2025)
por: Ding, Junchen, et al.
Publicado: (2025)
The Moral Machine Experiment on Large Language Models
por: Takemoto, Kazuhiro
Publicado: (2023)
por: Takemoto, Kazuhiro
Publicado: (2023)
Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations
por: Nunes, José Luiz, et al.
Publicado: (2024)
por: Nunes, José Luiz, et al.
Publicado: (2024)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
por: Backmann, Steffen, et al.
Publicado: (2025)
por: Backmann, Steffen, et al.
Publicado: (2025)
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
por: Zhou, Jingyan, et al.
Publicado: (2023)
por: Zhou, Jingyan, et al.
Publicado: (2023)
Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space
por: Lee, Minji, et al.
Publicado: (2024)
por: Lee, Minji, et al.
Publicado: (2024)
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
por: Wu, Ya, et al.
Publicado: (2025)
por: Wu, Ya, et al.
Publicado: (2025)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
por: Vida, Karina, et al.
Publicado: (2024)
por: Vida, Karina, et al.
Publicado: (2024)
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
por: Ridoy, Shahriyar Zaman, et al.
Publicado: (2025)
por: Ridoy, Shahriyar Zaman, et al.
Publicado: (2025)
Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language Models
por: Jung, Chani, et al.
Publicado: (2024)
por: Jung, Chani, et al.
Publicado: (2024)
MOKA: Moral Knowledge Augmentation for Moral Event Extraction
por: Zhang, Xinliang Frederick, et al.
Publicado: (2023)
por: Zhang, Xinliang Frederick, et al.
Publicado: (2023)
Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
por: Park, Sungwon, et al.
Publicado: (2024)
por: Park, Sungwon, et al.
Publicado: (2024)
MoralBench: Moral Evaluation of LLMs
por: Ji, Jianchao, et al.
Publicado: (2024)
por: Ji, Jianchao, et al.
Publicado: (2024)
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions
por: Preniqi, Vjosa, et al.
Publicado: (2024)
por: Preniqi, Vjosa, et al.
Publicado: (2024)
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
por: Morlat, Geoffroy, et al.
Publicado: (2025)
por: Morlat, Geoffroy, et al.
Publicado: (2025)
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
por: Liu, Jiarui, et al.
Publicado: (2025)
por: Liu, Jiarui, et al.
Publicado: (2025)
Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment
por: Hwang, Seojin, et al.
Publicado: (2026)
por: Hwang, Seojin, et al.
Publicado: (2026)
MoVa: Towards Generalizable Classification of Human Morals and Values
por: Chen, Ziyu, et al.
Publicado: (2025)
por: Chen, Ziyu, et al.
Publicado: (2025)
One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
por: Farid, Sualeha, et al.
Publicado: (2025)
por: Farid, Sualeha, et al.
Publicado: (2025)
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
por: Lamparth, Max, et al.
Publicado: (2024)
por: Lamparth, Max, et al.
Publicado: (2024)
Jailbreaking Large Language Models with Morality Attacks
por: Su, Ying, et al.
Publicado: (2026)
por: Su, Ying, et al.
Publicado: (2026)
Are Language Models Consequentialist or Deontological Moral Reasoners?
por: Samway, Keenan, et al.
Publicado: (2025)
por: Samway, Keenan, et al.
Publicado: (2025)
Literary Narrative as Moral Probe : A Cross-System Framework for Evaluating AI Ethical Reasoning and Refusal Behavior
por: Flynn, David C.
Publicado: (2026)
por: Flynn, David C.
Publicado: (2026)
Are Language Models Sensitive to Morally Irrelevant Distractors?
por: Shaw, Andrew, et al.
Publicado: (2026)
por: Shaw, Andrew, et al.
Publicado: (2026)
Histoires Morales: A French Dataset for Assessing Moral Alignment
por: Leteno, Thibaud, et al.
Publicado: (2025)
por: Leteno, Thibaud, et al.
Publicado: (2025)
Recent advances in interpretable machine learning using structure-based protein representations
por: Vecchietti, Luiz Felipe, et al.
Publicado: (2024)
por: Vecchietti, Luiz Felipe, et al.
Publicado: (2024)
Morally Programmed LLMs Reshape Human Morality
por: Lyu, Pengzhao, et al.
Publicado: (2026)
por: Lyu, Pengzhao, et al.
Publicado: (2026)
Ejemplares similares
-
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
por: Kim, Jiseon, et al.
Publicado: (2025) -
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
por: Kwon, Jea, et al.
Publicado: (2025) -
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
por: Kim, Minsung, et al.
Publicado: (2025) -
I Am Not Them: Fluid Identities and Persistent Out-group Bias in Large Language Models
por: Dong, Wenchao, et al.
Publicado: (2024) -
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
por: Park, Seongchan, et al.
Publicado: (2026)