Acting for the Right Reasons: Creating Reason-Sensitive Artificial Moral Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Baum, Kevin, Dargasz, Lisa, Jahn, Felix, Gros, Timo P., Wolf, Verena |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Integrating Reason-Based Moral Decision-Making in the Reinforcement Learning Architecture
by: Dargasz, Lisa
Published: (2025)
by: Dargasz, Lisa
Published: (2025)
Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment
by: Jahn, Felix, et al.
Published: (2026)
by: Jahn, Felix, et al.
Published: (2026)
Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance
by: Gabelmann, Julius, et al.
Published: (2026)
by: Gabelmann, Julius, et al.
Published: (2026)
GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
by: Yu, Jeffy, et al.
Published: (2024)
by: Yu, Jeffy, et al.
Published: (2024)
Inducing Human-like Biases in Moral Reasoning Language Models
by: Karpov, Artem, et al.
Published: (2024)
by: Karpov, Artem, et al.
Published: (2024)
Moral Alignment for LLM Agents
by: Tennant, Elizaveta, et al.
Published: (2024)
by: Tennant, Elizaveta, et al.
Published: (2024)
Dyna-Think: Synergizing Reasoning, Acting, and World Model Simulation in AI Agents
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
by: Sahoo, Subramanyam, et al.
Published: (2026)
by: Sahoo, Subramanyam, et al.
Published: (2026)
A New Paradigm for Counterfactual Reasoning in Fairness and Recourse
by: Bynum, Lucius E. J., et al.
Published: (2024)
by: Bynum, Lucius E. J., et al.
Published: (2024)
Artificial Intelligence Should Genuinely Support Clinical Reasoning and Decision Making To Bridge the Translational Gap
by: Sokol, Kacper, et al.
Published: (2025)
by: Sokol, Kacper, et al.
Published: (2025)
Interpretable Knowledge Tracing via Response Influence-based Counterfactual Reasoning
by: Cui, Jiajun, et al.
Published: (2023)
by: Cui, Jiajun, et al.
Published: (2023)
Dynamics of Moral Behavior in Heterogeneous Populations of Learning Agents
by: Tennant, Elizaveta, et al.
Published: (2024)
by: Tennant, Elizaveta, et al.
Published: (2024)
MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes
by: Chiu, Yu Ying, et al.
Published: (2025)
by: Chiu, Yu Ying, et al.
Published: (2025)
Holistically Evaluating the Environmental Impact of Creating Language Models
by: Morrison, Jacob, et al.
Published: (2025)
by: Morrison, Jacob, et al.
Published: (2025)
Creating a Cooperative AI Policymaking Platform through Open Source Collaboration
by: Lewington, Aiden, et al.
Published: (2024)
by: Lewington, Aiden, et al.
Published: (2024)
Building Interpretable Models for Moral Decision-Making
by: Goel, Mayank, et al.
Published: (2026)
by: Goel, Mayank, et al.
Published: (2026)
Leveraging Pedagogical Theories to Understand Student Learning Process with Graph-based Reasonable Knowledge Tracing
by: Cui, Jiajun, et al.
Published: (2024)
by: Cui, Jiajun, et al.
Published: (2024)
Deliberative Alignment: Reasoning Enables Safer Language Models
by: Guan, Melody Y., et al.
Published: (2024)
by: Guan, Melody Y., et al.
Published: (2024)
Reasoning-Based Refinement of Unsupervised Text Clusters with LLMs
by: Islam, Tunazzina
Published: (2026)
by: Islam, Tunazzina
Published: (2026)
Hybrid Approaches for Moral Value Alignment in AI Agents: a Manifesto
by: Tennant, Elizaveta, et al.
Published: (2023)
by: Tennant, Elizaveta, et al.
Published: (2023)
Selecting the Right LLM for eGov Explanations
by: Limonad, Lior, et al.
Published: (2025)
by: Limonad, Lior, et al.
Published: (2025)
Per-Domain Generalizing Policies: On Validation Instances and Scaling Behavior
by: Gros, Timo P., et al.
Published: (2025)
by: Gros, Timo P., et al.
Published: (2025)
When Reasoning Models Hurt Behavioral Simulation: A Solver-Sampler Mismatch in Multi-Agent LLM Negotiation
by: Andric, Sandro
Published: (2026)
by: Andric, Sandro
Published: (2026)
Personalized Decision Modeling: Utility Optimization or Textualized-Symbolic Reasoning
by: Zhao, Yibo, et al.
Published: (2025)
by: Zhao, Yibo, et al.
Published: (2025)
Benchmarking the Legal Reasoning of LLMs in Arabic Islamic Inheritance Cases
by: AlDahoul, Nouar, et al.
Published: (2025)
by: AlDahoul, Nouar, et al.
Published: (2025)
Contesting Artificial Moral Agents
by: Aijaz, Aisha
Published: (2026)
by: Aijaz, Aisha
Published: (2026)
Are There Exceptions to Goodhart's Law? On the Moral Justification of Fairness-Aware Machine Learning
by: Weerts, Hilde, et al.
Published: (2022)
by: Weerts, Hilde, et al.
Published: (2022)
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
by: Zhou, Andy, et al.
Published: (2023)
by: Zhou, Andy, et al.
Published: (2023)
Evaluating Reasoning LLMs for Suicide Screening with the Columbia-Suicide Severity Rating Scale
by: Patil, Avinash, et al.
Published: (2025)
by: Patil, Avinash, et al.
Published: (2025)
Operationalizing the Blueprint for an AI Bill of Rights: Recommendations for Practitioners, Researchers, and Policy Makers
by: Oesterling, Alex, et al.
Published: (2024)
by: Oesterling, Alex, et al.
Published: (2024)
A Theory of Inference Compute Scaling: Reasoning through Directed Stochastic Skill Search
by: Ellis-Mohr, Austin R., et al.
Published: (2025)
by: Ellis-Mohr, Austin R., et al.
Published: (2025)
The Mirage of Artificial Intelligence Terms of Use Restrictions
by: Henderson, Peter, et al.
Published: (2024)
by: Henderson, Peter, et al.
Published: (2024)
Artificial Intelligence Ecosystem for Automating Self-Directed Teaching
by: Gotavade, Tejas Satish
Published: (2024)
by: Gotavade, Tejas Satish
Published: (2024)
A Trustworthiness-based Metaphysics of Artificial Intelligence Systems
by: Ferrario, Andrea
Published: (2025)
by: Ferrario, Andrea
Published: (2025)
Responsible Artificial Intelligence: A Structured Literature Review
by: Goellner, Sabrina, et al.
Published: (2024)
by: Goellner, Sabrina, et al.
Published: (2024)
Automatic Evaluation Metrics for Artificially Generated Scientific Research
by: Höpner, Niklas, et al.
Published: (2025)
by: Höpner, Niklas, et al.
Published: (2025)
The Pursuit of Fairness in Artificial Intelligence Models: A Survey
by: Kheya, Tahsin Alamgir, et al.
Published: (2024)
by: Kheya, Tahsin Alamgir, et al.
Published: (2024)
Learning Fair Models without Sensitive Attributes: A Generative Approach
by: Zhu, Huaisheng, et al.
Published: (2022)
by: Zhu, Huaisheng, et al.
Published: (2022)
HH4AI: A methodological Framework for AI Human Rights impact assessment under the EUAI ACT
by: Ceravolo, Paolo, et al.
Published: (2025)
by: Ceravolo, Paolo, et al.
Published: (2025)
Generative Artificial Intelligence in Healthcare: Ethical Considerations and Assessment Checklist
by: Ning, Yilin, et al.
Published: (2023)
by: Ning, Yilin, et al.
Published: (2023)
Similar Items
-
Integrating Reason-Based Moral Decision-Making in the Reinforcement Learning Architecture
by: Dargasz, Lisa
Published: (2025) -
Breaking Up with Normatively Monolithic Agency with GRACE: A Reason-Based Neuro-Symbolic Architecture for Safe and Ethical AI Alignment
by: Jahn, Felix, et al.
Published: (2026) -
Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance
by: Gabelmann, Julius, et al.
Published: (2026) -
GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
by: Yu, Jeffy, et al.
Published: (2024) -
Inducing Human-like Biases in Moral Reasoning Language Models
by: Karpov, Artem, et al.
Published: (2024)