ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Thomas, Rohan Subramanian, Shiromani, Shikhar, Chaudhry, Abdullah, Li, Ruizhe, Sharma, Vasu, Zhu, Kevin, Dev, Sunishchal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
von: Sahay, Kenji, et al.
Veröffentlicht: (2025)
von: Sahay, Kenji, et al.
Veröffentlicht: (2025)
MoralBench: Moral Evaluation of LLMs
von: Ji, Jianchao, et al.
Veröffentlicht: (2024)
von: Ji, Jianchao, et al.
Veröffentlicht: (2024)
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
von: Shiromani, Shikhar, et al.
Veröffentlicht: (2026)
von: Shiromani, Shikhar, et al.
Veröffentlicht: (2026)
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
von: Agrawal, Shriyansh, et al.
Veröffentlicht: (2025)
von: Agrawal, Shriyansh, et al.
Veröffentlicht: (2025)
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
von: Batra, Shourya, et al.
Veröffentlicht: (2025)
von: Batra, Shourya, et al.
Veröffentlicht: (2025)
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
von: Rana, Manik, et al.
Veröffentlicht: (2025)
von: Rana, Manik, et al.
Veröffentlicht: (2025)
Ethical Reasoning and Moral Value Alignment of LLMs Depend on the Language we Prompt them in
von: Agarwal, Utkarsh, et al.
Veröffentlicht: (2024)
von: Agarwal, Utkarsh, et al.
Veröffentlicht: (2024)
Rethinking Machine Ethics -- Can LLMs Perform Moral Reasoning through the Lens of Moral Theories?
von: Zhou, Jingyan, et al.
Veröffentlicht: (2023)
von: Zhou, Jingyan, et al.
Veröffentlicht: (2023)
BengaliMoralBench: A Benchmark for Auditing Moral Reasoning in Large Language Models within Bengali Language and Culture
von: Ridoy, Shahriyar Zaman, et al.
Veröffentlicht: (2025)
von: Ridoy, Shahriyar Zaman, et al.
Veröffentlicht: (2025)
Exploring the psychology of LLMs' Moral and Legal Reasoning
von: Almeida, Guilherme F. C. F., et al.
Veröffentlicht: (2023)
von: Almeida, Guilherme F. C. F., et al.
Veröffentlicht: (2023)
Evaluating Moral Beliefs across LLMs through a Pluralistic Framework
von: Liu, Xuelin, et al.
Veröffentlicht: (2024)
von: Liu, Xuelin, et al.
Veröffentlicht: (2024)
Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering
von: Mao, Nathan, et al.
Veröffentlicht: (2026)
von: Mao, Nathan, et al.
Veröffentlicht: (2026)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
von: Singh, Harman, et al.
Veröffentlicht: (2024)
von: Singh, Harman, et al.
Veröffentlicht: (2024)
Evaluating Gender Bias of LLMs in Making Morality Judgements
von: Bajaj, Divij, et al.
Veröffentlicht: (2024)
von: Bajaj, Divij, et al.
Veröffentlicht: (2024)
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
von: Chaturvedi, Isha, et al.
Veröffentlicht: (2025)
von: Chaturvedi, Isha, et al.
Veröffentlicht: (2025)
NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
von: Gupta, Abhay, et al.
Veröffentlicht: (2025)
Text Prompt Injection of Vision Language Models
von: Zhu, Ruizhe
Veröffentlicht: (2025)
von: Zhu, Ruizhe
Veröffentlicht: (2025)
Moral Reasoning Across Languages: The Critical Role of Low-Resource Languages in LLMs
von: Zhou, Huichi, et al.
Veröffentlicht: (2025)
von: Zhou, Huichi, et al.
Veröffentlicht: (2025)
Emergent Persuasion: Will LLMs Persuade Without Being Prompted?
von: Chang, Vincent, et al.
Veröffentlicht: (2025)
von: Chang, Vincent, et al.
Veröffentlicht: (2025)
Moral Lenses, Political Coordinates: Towards Ideological Positioning of Morally Conditioned LLMs
von: Yuan, Chenchen, et al.
Veröffentlicht: (2026)
von: Yuan, Chenchen, et al.
Veröffentlicht: (2026)
Procedural Dilemma Generation for Evaluating Moral Reasoning in Humans and Language Models
von: Fränken, Jan-Philipp, et al.
Veröffentlicht: (2024)
von: Fränken, Jan-Philipp, et al.
Veröffentlicht: (2024)
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
von: Chiu, Yu Ying, et al.
Veröffentlicht: (2025)
MFTCXplain: A Multilingual Benchmark Dataset for Evaluating the Moral Reasoning of LLMs through Multi-hop Hate Speech Explanation
von: Trager, Jackson, et al.
Veröffentlicht: (2025)
von: Trager, Jackson, et al.
Veröffentlicht: (2025)
Rosetta-PL: Propositional Logic as a Benchmark for Large Language Model Reasoning
von: Baek, Shaun, et al.
Veröffentlicht: (2025)
von: Baek, Shaun, et al.
Veröffentlicht: (2025)
Are Rules Meant to be Broken? Understanding Multilingual Moral Reasoning as a Computational Pipeline with UniMoral
von: Kumar, Shivani, et al.
Veröffentlicht: (2025)
von: Kumar, Shivani, et al.
Veröffentlicht: (2025)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
von: He, Jiahang, et al.
Veröffentlicht: (2025)
von: He, Jiahang, et al.
Veröffentlicht: (2025)
Probe-Rewrite-Evaluate: A Workflow for Reliable Benchmarks and Quantifying Evaluation Awareness
von: Xiong, Lang, et al.
Veröffentlicht: (2025)
von: Xiong, Lang, et al.
Veröffentlicht: (2025)
One Model, Many Morals: Uncovering Cross-Linguistic Misalignments in Computational Moral Reasoning
von: Farid, Sualeha, et al.
Veröffentlicht: (2025)
von: Farid, Sualeha, et al.
Veröffentlicht: (2025)
Are Language Models Consequentialist or Deontological Moral Reasoners?
von: Samway, Keenan, et al.
Veröffentlicht: (2025)
von: Samway, Keenan, et al.
Veröffentlicht: (2025)
FairMT-Bench: Benchmarking Fairness for Multi-turn Dialogue in Conversational LLMs
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
von: Fan, Zhiting, et al.
Veröffentlicht: (2024)
That is Unacceptable: the Moral Foundations of Canceling
von: Lo, Soda Marem, et al.
Veröffentlicht: (2025)
von: Lo, Soda Marem, et al.
Veröffentlicht: (2025)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
von: Afonin, Nikita, et al.
Veröffentlicht: (2025)
von: Afonin, Nikita, et al.
Veröffentlicht: (2025)
Adaptive Originality Filtering: Rejection Based Prompting and RiddleScore for Culturally Grounded Multilingual Riddle Generation
von: Le, Duy, et al.
Veröffentlicht: (2025)
von: Le, Duy, et al.
Veröffentlicht: (2025)
Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs
von: Huang, Jen-tse, et al.
Veröffentlicht: (2026)
von: Huang, Jen-tse, et al.
Veröffentlicht: (2026)
ProSA: Assessing and Understanding the Prompt Sensitivity of LLMs
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024)
von: Zhuo, Jingming, et al.
Veröffentlicht: (2024)
MOKA: Moral Knowledge Augmentation for Moral Event Extraction
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2023)
von: Zhang, Xinliang Frederick, et al.
Veröffentlicht: (2023)
SG-Bench: Evaluating LLM Safety Generalization Across Diverse Tasks and Prompt Types
von: Mou, Yutao, et al.
Veröffentlicht: (2024)
von: Mou, Yutao, et al.
Veröffentlicht: (2024)
CA-BED: Conversation-Aware Bayesian Experimental Design
von: Arnould, Daniel, et al.
Veröffentlicht: (2026)
von: Arnould, Daniel, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
COMPASS: Context-Modulated PID Attention Steering System for Hallucination Mitigation
von: Sahay, Kenji, et al.
Veröffentlicht: (2025) -
MoralBench: Moral Evaluation of LLMs
von: Ji, Jianchao, et al.
Veröffentlicht: (2024) -
The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders
von: Shiromani, Shikhar, et al.
Veröffentlicht: (2026) -
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
von: Agrawal, Shriyansh, et al.
Veröffentlicht: (2025) -
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
von: Batra, Shourya, et al.
Veröffentlicht: (2025)