Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Bocheng, Chen, Xi, Zi, Han, Mao, Haitao, Qi, Zimo, Zhang, Xitong, Johnson, Kristen, Liu, Guangliang
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917532664856576
author Chen, Bocheng
Chen, Xi
Zi, Han
Mao, Haitao
Qi, Zimo
Zhang, Xitong
Johnson, Kristen
Liu, Guangliang
author_facet Chen, Bocheng
Chen, Xi
Zi, Han
Mao, Haitao
Qi, Zimo
Zhang, Xitong
Johnson, Kristen
Liu, Guangliang
contents Moral sensitivity is the most fundamental capability underlying human moral competence. Although many approaches aim to align large language models (LLMs) with human moral values, they primarily focus on fitting the distributions of morally appropriate texts while overlooking how to enable moral sensitivity acquisition in LLMs. In this paper, we take a step toward addressing the question: How can moral sensitivity be acquired in LLMs? Specifically, we propose a pragmatic inference approach that facilitates moral sensitivity acquisition in LLMs by enabling them to diagnose and correct moral errors. A central strength of our pragmatic inference approach lies in its unified perspective: rather than modeling moral discourses across semantically diverse and complex surface forms, it provides a principled framework for designing pragmatic inference procedures grounded in their inferential load. Empirical evidence demonstrates that our pragmatic approach can enable moral sensitivity acquisition in LLMs and generalizes effectively across tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2601_03079
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models
Chen, Bocheng
Chen, Xi
Zi, Han
Mao, Haitao
Qi, Zimo
Zhang, Xitong
Johnson, Kristen
Liu, Guangliang
Computation and Language
Moral sensitivity is the most fundamental capability underlying human moral competence. Although many approaches aim to align large language models (LLMs) with human moral values, they primarily focus on fitting the distributions of morally appropriate texts while overlooking how to enable moral sensitivity acquisition in LLMs. In this paper, we take a step toward addressing the question: How can moral sensitivity be acquired in LLMs? Specifically, we propose a pragmatic inference approach that facilitates moral sensitivity acquisition in LLMs by enabling them to diagnose and correct moral errors. A central strength of our pragmatic inference approach lies in its unified perspective: rather than modeling moral discourses across semantically diverse and complex surface forms, it provides a principled framework for designing pragmatic inference procedures grounded in their inferential load. Empirical evidence demonstrates that our pragmatic approach can enable moral sensitivity acquisition in LLMs and generalizes effectively across tasks.
title Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models
topic Computation and Language
url https://arxiv.org/abs/2601.03079