MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
Fuente:
arXiv
Saved in:
| Main Authors: | Chiu, Yu Ying, Lee, Michael S., Calcott, Rachel, Handoko, Brandon, de Font-Reaulx, Paul, Rodriguez, Paula, Zhang, Chen Bo Calvin, Han, Ziwen, Sehwag, Udari Madhushani, Maurya, Yash, Knight, Christina Q, Lloyd, Harry R., Bacus, Florence, Mazeika, Mantas, Liu, Bing, Choi, Yejin, Gordon, Mitchell L, Levine, Sydney |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
In-Context Learning with Topological Information for Knowledge Graph Completion
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
by: Sehwag, Udari Madhushani, et al.
Published: (2025)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
LHAW: Controllable Underspecification for Long-Horizon Tasks
by: Pu, George, et al.
Published: (2026)
by: Pu, George, et al.
Published: (2026)
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
by: Lee, Michael S., et al.
Published: (2026)
by: Lee, Michael S., et al.
Published: (2026)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
by: Cook, Thomas, et al.
Published: (2025)
by: Cook, Thomas, et al.
Published: (2025)
TextQuests: How Good are LLMs at Text-Based Video Games?
by: Phan, Long, et al.
Published: (2025)
by: Phan, Long, et al.
Published: (2025)
DeliberationBench: A Normative Benchmark for the Influence of Large Language Models on Users' Views
by: Hewitt, Luke, et al.
Published: (2026)
by: Hewitt, Luke, et al.
Published: (2026)
Intuitions of Compromise: Utilitarianism vs. Contractualism
by: Moore, Jared, et al.
Published: (2024)
by: Moore, Jared, et al.
Published: (2024)
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
by: Sorensen, Taylor, et al.
Published: (2023)
by: Sorensen, Taylor, et al.
Published: (2023)
Can Language Models Reason about Individualistic Human Values and Preferences?
by: Jiang, Liwei, et al.
Published: (2024)
by: Jiang, Liwei, et al.
Published: (2024)
Collab: Controlled Decoding using Mixture of Agents for LLM Alignment
by: Chakraborty, Souradip, et al.
Published: (2025)
by: Chakraborty, Souradip, et al.
Published: (2025)
O3D: Offline Data-driven Discovery and Distillation for Sequential Decision-Making with Large Language Models
by: Xiao, Yuchen, et al.
Published: (2023)
by: Xiao, Yuchen, et al.
Published: (2023)
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
by: Feng, Shangbin, et al.
Published: (2024)
by: Feng, Shangbin, et al.
Published: (2024)
AI Governance and Accountability: An Analysis of Anthropic's Claude
by: Priyanshu, Aman, et al.
Published: (2024)
by: Priyanshu, Aman, et al.
Published: (2024)
A Roadmap to Pluralistic Alignment
by: Sorensen, Taylor, et al.
Published: (2024)
by: Sorensen, Taylor, et al.
Published: (2024)
AdvBDGen: Adversarially Fortified Prompt-Specific Fuzzy Backdoor Generator Against LLM Alignment
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
by: Pathmanathan, Pankayaraj, et al.
Published: (2024)
Can LLMs be Scammed? A Baseline Measurement Study
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
by: Sehwag, Udari Madhushani, et al.
Published: (2024)
Pluralistic Leaderboards
by: Haghtalab, Nika, et al.
Published: (2026)
by: Haghtalab, Nika, et al.
Published: (2026)
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
by: Mazeika, Mantas, et al.
Published: (2025)
by: Mazeika, Mantas, et al.
Published: (2025)
Machine Theory of Mind and the Structure of Human Values
by: de Font-Reaulx, Paul
Published: (2025)
by: de Font-Reaulx, Paul
Published: (2025)
Order-Based Pre-training Strategies for Procedural Text Understanding
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
by: Wei, Boyi, et al.
Published: (2025)
by: Wei, Boyi, et al.
Published: (2025)
Who Gets Flagged? The Pluralistic Evaluation Gap in AI Content Watermarking
by: Nemecek, Alexander, et al.
Published: (2026)
by: Nemecek, Alexander, et al.
Published: (2026)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
by: Mireshghallah, Niloofar, et al.
Published: (2024)
by: Mireshghallah, Niloofar, et al.
Published: (2024)
Pluralistic Salient Object Detection
by: Feng, Xuelu, et al.
Published: (2024)
by: Feng, Xuelu, et al.
Published: (2024)
SPICA: Retrieving Scenarios for Pluralistic In-Context Alignment
by: Chen, Quan Ze, et al.
Published: (2024)
by: Chen, Quan Ze, et al.
Published: (2024)
PERSONA: A Reproducible Testbed for Pluralistic Alignment
by: Castricato, Louis, et al.
Published: (2024)
by: Castricato, Louis, et al.
Published: (2024)
Unified Locational Differential Privacy Framework
by: Priyanshu, Aman, et al.
Published: (2024)
by: Priyanshu, Aman, et al.
Published: (2024)
ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans
by: Sadana, Ananya, et al.
Published: (2025)
by: Sadana, Ananya, et al.
Published: (2025)
Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
by: Ren, Richard, et al.
Published: (2024)
by: Ren, Richard, et al.
Published: (2024)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
by: Li, Jing-Jing, et al.
Published: (2024)
by: Li, Jing-Jing, et al.
Published: (2024)
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
by: Kachwala, Zoher, et al.
Published: (2026)
by: Kachwala, Zoher, et al.
Published: (2026)
ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
by: Sehwag, Udari Madhushani, et al.
Published: (2026)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
by: Mazeika, Mantas, et al.
Published: (2024)
by: Mazeika, Mantas, et al.
Published: (2024)
Overton Pluralistic Reinforcement Learning for Large Language Models
by: Fu, Yu, et al.
Published: (2026)
by: Fu, Yu, et al.
Published: (2026)
Plan-and-Write: Structure-Guided Length Control for LLMs without Model Retraining
by: Akinfaderin, Adewale, et al.
Published: (2025)
by: Akinfaderin, Adewale, et al.
Published: (2025)
Pluralistic Off-policy Evaluation and Alignment
by: Huang, Chengkai, et al.
Published: (2025)
by: Huang, Chengkai, et al.
Published: (2025)
DailyDilemmas: Revealing Value Preferences of LLMs with Quandaries of Daily Life
by: Chiu, Yu Ying, et al.
Published: (2024)
by: Chiu, Yu Ying, et al.
Published: (2024)
Similar Items
-
In-Context Learning with Topological Information for Knowledge Graph Completion
by: Sehwag, Udari Madhushani, et al.
Published: (2024) -
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
by: Sehwag, Udari Madhushani, et al.
Published: (2025) -
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
by: Xu, Yuancheng, et al.
Published: (2024) -
LHAW: Controllable Underspecification for Long-Horizon Tasks
by: Pu, George, et al.
Published: (2026) -
ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
by: Lee, Michael S., et al.
Published: (2026)