Exploring Persona-dependent LLM Alignment for the Moral Machine Experiment
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Jiseon, Kwon, Jea, Vecchietti, Luiz Felipe, Oh, Alice, Cha, Meeyoung |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
por: Kwon, Jea, et al.
Publicado: (2025)
por: Kwon, Jea, et al.
Publicado: (2025)
Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions
por: Kim, Jiseon, et al.
Publicado: (2026)
por: Kim, Jiseon, et al.
Publicado: (2026)
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
por: Kim, Minsung, et al.
Publicado: (2025)
por: Kim, Minsung, et al.
Publicado: (2025)
Uncovering Factor Level Preferences to Improve Human-Model Alignment
por: Oh, Juhyun, et al.
Publicado: (2024)
por: Oh, Juhyun, et al.
Publicado: (2024)
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
por: Vida, Karina, et al.
Publicado: (2024)
por: Vida, Karina, et al.
Publicado: (2024)
Moral Susceptibility and Robustness under Persona Role-Play in Large Language Models
por: Costa, Davi Bastos, et al.
Publicado: (2025)
por: Costa, Davi Bastos, et al.
Publicado: (2025)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
por: Hu, Yueqing, et al.
Publicado: (2026)
por: Hu, Yueqing, et al.
Publicado: (2026)
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
por: Kiet, Huynh Trung, et al.
Publicado: (2026)
por: Kiet, Huynh Trung, et al.
Publicado: (2026)
German General Social Survey Personas: A Survey-Derived Persona Prompt Collection for Population-Aligned LLM Studies
por: Rupprecht, Jens, et al.
Publicado: (2025)
por: Rupprecht, Jens, et al.
Publicado: (2025)
RoleConflictBench: A Benchmark of Role Conflict Scenarios for Evaluating LLMs' Contextual Sensitivity
por: Shin, Jisu, et al.
Publicado: (2025)
por: Shin, Jisu, et al.
Publicado: (2025)
Societal Alignment Frameworks Can Improve LLM Alignment
por: Stańczak, Karolina, et al.
Publicado: (2025)
por: Stańczak, Karolina, et al.
Publicado: (2025)
Scaling Law in LLM Simulated Personality: More Detailed and Realistic Persona Profile Is All You Need
por: Bai, Yuqi, et al.
Publicado: (2025)
por: Bai, Yuqi, et al.
Publicado: (2025)
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas
por: Backmann, Steffen, et al.
Publicado: (2025)
por: Backmann, Steffen, et al.
Publicado: (2025)
KoBBQ: Korean Bias Benchmark for Question Answering
por: Jin, Jiho, et al.
Publicado: (2023)
por: Jin, Jiho, et al.
Publicado: (2023)
Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks
por: Greco, Candida M., et al.
Publicado: (2026)
por: Greco, Candida M., et al.
Publicado: (2026)
Does Cross-Cultural Alignment Change the Commonsense Morality of Language Models?
por: Jinnai, Yuu
Publicado: (2024)
por: Jinnai, Yuu
Publicado: (2024)
The Generative AI Paradox on Evaluation: What It Can Solve, It May Not Evaluate
por: Oh, Juhyun, et al.
Publicado: (2024)
por: Oh, Juhyun, et al.
Publicado: (2024)
From Descriptive to Prescriptive: Uncover the Social Value Alignment of LLM-based Agents
por: Qu, Jinxian, et al.
Publicado: (2026)
por: Qu, Jinxian, et al.
Publicado: (2026)
PerMix-RLVR: Preserving Persona Expressivity under Verifiable-Reward Alignment
por: Oh, Jihwan, et al.
Publicado: (2026)
por: Oh, Jihwan, et al.
Publicado: (2026)
LLM Generated Persona is a Promise with a Catch
por: Li, Ang, et al.
Publicado: (2025)
por: Li, Ang, et al.
Publicado: (2025)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
por: Bojic, Ljubisa, et al.
Publicado: (2026)
por: Bojic, Ljubisa, et al.
Publicado: (2026)
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
por: Oh, Gyutaek, et al.
Publicado: (2025)
por: Oh, Gyutaek, et al.
Publicado: (2025)
Machine Learning for Detection and Analysis of Novel LLM Jailbreaks
por: Hawkins, John, et al.
Publicado: (2025)
por: Hawkins, John, et al.
Publicado: (2025)
Synthetic Reader Panels: Tournament-Based Ideation with LLM Personas for Autonomous Publishing
por: Zimmerman, Fred
Publicado: (2026)
por: Zimmerman, Fred
Publicado: (2026)
Empirical Evidence for Alignment Faking in a Small LLM and Prompt-Based Mitigation Techniques
por: Koorndijk, Jeanice
Publicado: (2025)
por: Koorndijk, Jeanice
Publicado: (2025)
The Need for a Socially-Grounded Persona Framework for User Simulation
por: Venkit, Pranav Narayanan, et al.
Publicado: (2026)
por: Venkit, Pranav Narayanan, et al.
Publicado: (2026)
Steering at the Source: Style Modulation Heads for Robust Persona Control
por: Izawa, Yoshihiro, et al.
Publicado: (2026)
por: Izawa, Yoshihiro, et al.
Publicado: (2026)
Moral Mazes in the Era of LLMs
por: Nguyen, Dang, et al.
Publicado: (2026)
por: Nguyen, Dang, et al.
Publicado: (2026)
Between Rules and Reality: On the Context Sensitivity of LLM Moral Judgment
por: Sauter, Adrian, et al.
Publicado: (2026)
por: Sauter, Adrian, et al.
Publicado: (2026)
A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
por: Venkit, Pranav Narayanan, et al.
Publicado: (2025)
por: Venkit, Pranav Narayanan, et al.
Publicado: (2025)
ProgressGym: Alignment with a Millennium of Moral Progress
por: Qiu, Tianyi, et al.
Publicado: (2024)
por: Qiu, Tianyi, et al.
Publicado: (2024)
Spotting Out-of-Character Behavior: Atomic-Level Evaluation of Persona Fidelity in Open-Ended Generation
por: Shin, Jisu, et al.
Publicado: (2025)
por: Shin, Jisu, et al.
Publicado: (2025)
Scopes of Alignment
por: Varshney, Kush R., et al.
Publicado: (2025)
por: Varshney, Kush R., et al.
Publicado: (2025)
EconCausal: A Context-Aware Economic Reasoning Benchmark for Large Language Models
por: Lee, Donggyu, et al.
Publicado: (2025)
por: Lee, Donggyu, et al.
Publicado: (2025)
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
por: Biancotti, Claudia, et al.
Publicado: (2024)
por: Biancotti, Claudia, et al.
Publicado: (2024)
Moral Alignment for LLM Agents
por: Tennant, Elizaveta, et al.
Publicado: (2024)
por: Tennant, Elizaveta, et al.
Publicado: (2024)
Widespread Gender and Pronoun Bias in Moral Judgments Across LLMs
por: Fernandes, Gustavo Lúcius, et al.
Publicado: (2026)
por: Fernandes, Gustavo Lúcius, et al.
Publicado: (2026)
Moral Outrage Shapes Commitments Beyond Attention: Multimodal Moral Emotions on YouTube in Korea and the US
por: Park, Seongchan, et al.
Publicado: (2026)
por: Park, Seongchan, et al.
Publicado: (2026)
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
por: Tzachristas, Ioannis, et al.
Publicado: (2025)
por: Tzachristas, Ioannis, et al.
Publicado: (2025)
DetectAnyLLM: Towards Generalizable and Robust Detection of Machine-Generated Text Across Domains and Models
por: Fu, Jiachen, et al.
Publicado: (2025)
por: Fu, Jiachen, et al.
Publicado: (2025)
Ejemplares similares
-
Dropouts in Confidence: Moral Uncertainty in Human-LLM Alignment
por: Kwon, Jea, et al.
Publicado: (2025) -
Machine Behavior in Relational Moral Dilemmas: Moral Rightness, Predicted Human Behavior, and Model Decisions
por: Kim, Jiseon, et al.
Publicado: (2026) -
How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
por: Kim, Minsung, et al.
Publicado: (2025) -
Uncovering Factor Level Preferences to Improve Human-Model Alignment
por: Oh, Juhyun, et al.
Publicado: (2024) -
Decoding Multilingual Moral Preferences: Unveiling LLM's Biases Through the Moral Machine Experiment
por: Vida, Karina, et al.
Publicado: (2024)