Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Jiseon, Kwon, Jea, Vecchietti, Luiz Felipe, Dong, Wenchao, Kim, Jaehong, Cha, Meeyoung |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
di: Kim, Jiseon, et al.
Pubblicazione: (2025)
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
di: Kwon, Jea, et al.
Pubblicazione: (2025)
di: Kwon, Jea, et al.
Pubblicazione: (2025)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
di: Kim, Minsung, et al.
Pubblicazione: (2025)
di: Kim, Minsung, et al.
Pubblicazione: (2025)
I Am Not Them: Fluid Identities and Persistent Out-group Bias in Large Language Models
di: Dong, Wenchao, et al.
Pubblicazione: (2024)
di: Dong, Wenchao, et al.
Pubblicazione: (2024)
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
di: Park, Seongchan, et al.
Pubblicazione: (2026)
di: Park, Seongchan, et al.
Pubblicazione: (2026)
Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models
di: Fränken, Jan-Philipp, et al.
Pubblicazione: (2024)
di: Fränken, Jan-Philipp, et al.
Pubblicazione: (2024)
Persona Setting Pitfall: Persistent Outgroup Biases in Large Language Models Arising from Social Identity Adoption
di: Dong, Wenchao, et al.
Pubblicazione: (2024)
di: Dong, Wenchao, et al.
Pubblicazione: (2024)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
di: Oh, Juhyun, et al.
Pubblicazione: (2024)
di: Oh, Juhyun, et al.
Pubblicazione: (2024)
Bilinear representation mitigates reversal curse and enables consistent model editing
di: Kim, Dong-Kyum, et al.
Pubblicazione: (2025)
di: Kim, Dong-Kyum, et al.
Pubblicazione: (2025)
The Greatest Good Benchmark: Measuring LLMs' Alignment with Utilitarian Moral Dilemmas
di: Marraffini, Giovanni Franco Gabriel, et al.
Pubblicazione: (2025)
di: Marraffini, Giovanni Franco Gabriel, et al.
Pubblicazione: (2025)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
di: Lee, Donggyu, et al.
Pubblicazione: (2025)
di: Lee, Donggyu, et al.
Pubblicazione: (2025)
Erase or Hide? Suppressing Spurious Unlearning Neurons for Robust Unlearning
di: Yang, Nakyeong, et al.
Pubblicazione: (2025)
di: Yang, Nakyeong, et al.
Pubblicazione: (2025)
How You Ask Matters! Adaptive RAG Robustness to Query Variations
di: Jang, Yunah, et al.
Pubblicazione: (2026)
di: Jang, Yunah, et al.
Pubblicazione: (2026)
"Pull or Not to Pull?'': Investigating Moral Biases in Leading Large Language Models Across Ethical Dilemmas
di: Ding, Junchen, et al.
Pubblicazione: (2025)
di: Ding, Junchen, et al.
Pubblicazione: (2025)
The Moral Machine Experiment on Large Language Models
di: Takemoto, Kazuhiro
Pubblicazione: (2023)
di: Takemoto, Kazuhiro
Pubblicazione: (2023)
Are Large Language Models Moral Hypocrites? A Study Based on Moral Foundations
di: Nunes, José Luiz, et al.
Pubblicazione: (2024)
di: Nunes, José Luiz, et al.
Pubblicazione: (2024)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
di: Backmann, Steffen, et al.
Pubblicazione: (2025)
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
di: Zhou, Jingyan, et al.
Pubblicazione: (2023)
di: Zhou, Jingyan, et al.
Pubblicazione: (2023)
Robust Optimization in Protein Fitness Landscapes Using Reinforcement Learning in Latent Space
di: Lee, Minji, et al.
Pubblicazione: (2024)
di: Lee, Minji, et al.
Pubblicazione: (2024)
The Staircase of Ethics: Probing LLM Value Priorities through Multi-Step Induction to Complex Moral Dilemmas
di: Wu, Ya, et al.
Pubblicazione: (2025)
di: Wu, Ya, et al.
Pubblicazione: (2025)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
di: Vida, Karina, et al.
Pubblicazione: (2024)
di: Vida, Karina, et al.
Pubblicazione: (2024)
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
di: Ridoy, Shahriyar Zaman, et al.
Pubblicazione: (2025)
di: Ridoy, Shahriyar Zaman, et al.
Pubblicazione: (2025)
Perceptions to Beliefs: Exploring Precursory Inferences for Theory of Mind in Large Language Models
di: Jung, Chani, et al.
Pubblicazione: (2024)
di: Jung, Chani, et al.
Pubblicazione: (2024)
MOKA: Moral Knowledge Augmentation for Moral Event Extraction
di: Zhang, Xinliang Frederick, et al.
Pubblicazione: (2023)
di: Zhang, Xinliang Frederick, et al.
Pubblicazione: (2023)
Adversarial Style Augmentation via Large Language Model for Robust Fake News Detection
di: Park, Sungwon, et al.
Pubblicazione: (2024)
di: Park, Sungwon, et al.
Pubblicazione: (2024)
MoralBench: Moral Evaluation of LLMs
di: Ji, Jianchao, et al.
Pubblicazione: (2024)
di: Ji, Jianchao, et al.
Pubblicazione: (2024)
MoralBERT: A Fine-Tuned Language Model for Capturing Moral Values in Social Discussions
di: Preniqi, Vjosa, et al.
Pubblicazione: (2024)
di: Preniqi, Vjosa, et al.
Pubblicazione: (2024)
Morality is Contextual: Learning Interpretable Moral Contexts from Human Data with Probabilistic Clustering and Large Language Models
di: Morlat, Geoffroy, et al.
Pubblicazione: (2025)
di: Morlat, Geoffroy, et al.
Pubblicazione: (2025)
Synthetic Socratic Debates: Examining Persona Effects on Moral Decision and Persuasion Dynamics
di: Liu, Jiarui, et al.
Pubblicazione: (2025)
di: Liu, Jiarui, et al.
Pubblicazione: (2025)
Framing Matters: Addressing Framing Sensitivity in Decision-Making through Behaviorally-Grounded Value Alignment
di: Hwang, Seojin, et al.
Pubblicazione: (2026)
di: Hwang, Seojin, et al.
Pubblicazione: (2026)
MoVa: Towards Generalizable Classification of Human Morals and Values
di: Chen, Ziyu, et al.
Pubblicazione: (2025)
di: Chen, Ziyu, et al.
Pubblicazione: (2025)
One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
di: Farid, Sualeha, et al.
Pubblicazione: (2025)
di: Farid, Sualeha, et al.
Pubblicazione: (2025)
Human vs. Machine: Behavioral Differences Between Expert Humans and Language Models in Wargame Simulations
di: Lamparth, Max, et al.
Pubblicazione: (2024)
di: Lamparth, Max, et al.
Pubblicazione: (2024)
Jailbreaking Large Language Models with Morality Attacks
di: Su, Ying, et al.
Pubblicazione: (2026)
di: Su, Ying, et al.
Pubblicazione: (2026)
Are Language Models Consequentialist or Deontological Moral Reasoners?
di: Samway, Keenan, et al.
Pubblicazione: (2025)
di: Samway, Keenan, et al.
Pubblicazione: (2025)
Literary Narrative as Moral Probe : A Cross-System Framework for Evaluating AI Ethical Reasoning and Refusal Behavior
di: Flynn, David C.
Pubblicazione: (2026)
di: Flynn, David C.
Pubblicazione: (2026)
Are Language Models Sensitive to Morally Irrelevant Distractors?
di: Shaw, Andrew, et al.
Pubblicazione: (2026)
di: Shaw, Andrew, et al.
Pubblicazione: (2026)
Histoires Morales: A French Dataset for Assessing Moral Alignment
di: Leteno, Thibaud, et al.
Pubblicazione: (2025)
di: Leteno, Thibaud, et al.
Pubblicazione: (2025)
Recent advances in interpretable machine learning using structure-based protein representations
di: Vecchietti, Luiz Felipe, et al.
Pubblicazione: (2024)
di: Vecchietti, Luiz Felipe, et al.
Pubblicazione: (2024)
Morally Programmed LLMs Reshape Human Morality
di: Lyu, Pengzhao, et al.
Pubblicazione: (2026)
di: Lyu, Pengzhao, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
di: Kim, Jiseon, et al.
Pubblicazione: (2025) -
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
di: Kwon, Jea, et al.
Pubblicazione: (2025) -
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
di: Kim, Minsung, et al.
Pubblicazione: (2025) -
I Am Not Them: Fluid Identities and Persistent Out-group Bias in Large Language Models
di: Dong, Wenchao, et al.
Pubblicazione: (2024) -
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
di: Park, Seongchan, et al.
Pubblicazione: (2026)