Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Bocheng, Chen, Xi, Zi, Han, Mao, Haitao, Qi, Zimo, Zhang, Xitong, Johnson, Kristen, Liu, Guangliang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalization
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Pragmatic Inference for Moral Reasoning Acquisition: Generalization via Metapragmatic Links
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Discourse Heuristics For Paradoxically Moral Self-Correction
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
On the Convergence of Moral Self-Correction in Large Language Models
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Smaller Large Language Models Can Do Moral Self-Correction
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
Self-correction is Not An Innate Capability in Language Models
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
On the Intrinsic Self-Correction Capability of LLMs: Uncertainty and Latent Concept
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
A Survey to Recent Progress Towards Understanding In-Context Learning
von: Mao, Haitao, et al.
Veröffentlicht: (2024)
von: Mao, Haitao, et al.
Veröffentlicht: (2024)
Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay
von: Yusoff, Mariah Al Giptiah Binte, et al.
Veröffentlicht: (2026)
von: Yusoff, Mariah Al Giptiah Binte, et al.
Veröffentlicht: (2026)
Towards Understanding Task-agnostic Debiasing Through the Lenses of Intrinsic Bias and Forgetfulness
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
Deactivating Refusal Triggers: Understanding and Mitigating Overrefusal in Safety Alignment
von: Xue, Zhiyu, et al.
Veröffentlicht: (2026)
von: Xue, Zhiyu, et al.
Veröffentlicht: (2026)
No Free Lunch for Defending Against Prefilling Attack by In-Context Learning
von: Xue, Zhiyu, et al.
Veröffentlicht: (2024)
von: Xue, Zhiyu, et al.
Veröffentlicht: (2024)
Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework
von: Xu, Zishan, et al.
Veröffentlicht: (2025)
von: Xu, Zishan, et al.
Veröffentlicht: (2025)
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
von: Huang, Shiting, et al.
Veröffentlicht: (2025)
von: Huang, Shiting, et al.
Veröffentlicht: (2025)
Language Acquisition Device in Large Language Models
von: Mita, Masato, et al.
Veröffentlicht: (2026)
von: Mita, Masato, et al.
Veröffentlicht: (2026)
Large Language Models Are State-of-the-Art Evaluator for Grammatical Error Correction
von: Kobayashi, Masamune, et al.
Veröffentlicht: (2024)
von: Kobayashi, Masamune, et al.
Veröffentlicht: (2024)
Rethinking the Roles of Large Language Models in Chinese Grammatical Error Correction
von: Li, Yinghui, et al.
Veröffentlicht: (2024)
von: Li, Yinghui, et al.
Veröffentlicht: (2024)
ASR Error Correction using Large Language Models
von: Ma, Rao, et al.
Veröffentlicht: (2024)
von: Ma, Rao, et al.
Veröffentlicht: (2024)
Are Language Models Sensitive to Morally Irrelevant Distractors?
von: Shaw, Andrew, et al.
Veröffentlicht: (2026)
von: Shaw, Andrew, et al.
Veröffentlicht: (2026)
BLADE: Enhancing Black-box Large Language Models with Small Domain-Specific Models
von: Li, Haitao, et al.
Veröffentlicht: (2024)
von: Li, Haitao, et al.
Veröffentlicht: (2024)
Label-free Node Classification on Graphs with Large Language Models (LLMS)
von: Chen, Zhikai, et al.
Veröffentlicht: (2023)
von: Chen, Zhikai, et al.
Veröffentlicht: (2023)
Role and Relevance of the Learners’ Errors in Second Language Acquisition
von: Dr. Vinay Kumar Singh
Veröffentlicht: (2017)
von: Dr. Vinay Kumar Singh
Veröffentlicht: (2017)
SMRC: Aligning Large Language Models with Student Reasoning for Mathematical Error Correction
von: Zeng, Biaojie, et al.
Veröffentlicht: (2025)
von: Zeng, Biaojie, et al.
Veröffentlicht: (2025)
Towards Robust Instruction Tuning on Multimodal Large Language Models
von: Han, Wei, et al.
Veröffentlicht: (2024)
von: Han, Wei, et al.
Veröffentlicht: (2024)
Computational Reasoning of Large Language Models
von: Wu, Haitao, et al.
Veröffentlicht: (2025)
von: Wu, Haitao, et al.
Veröffentlicht: (2025)
The Dark Side of Human Feedback: Poisoning Large Language Models via User Inputs
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
von: Chen, Bocheng, et al.
Veröffentlicht: (2024)
Large Language Models based ASR Error Correction for Child Conversations
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
von: Xu, Anfeng, et al.
Veröffentlicht: (2025)
A Language-agnostic Model of Child Language Acquisition
von: Mahon, Louis, et al.
Veröffentlicht: (2024)
von: Mahon, Louis, et al.
Veröffentlicht: (2024)
Prompting Large Language Models with Human Error Markings for Self-Correcting Machine Translation
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)
von: Berger, Nathaniel, et al.
Veröffentlicht: (2024)
Evaluating the Capability of Large-scale Language Models on Chinese Grammatical Error Correction Task
von: Qu, Fanyi, et al.
Veröffentlicht: (2023)
von: Qu, Fanyi, et al.
Veröffentlicht: (2023)
Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
von: Xia, Han, et al.
Veröffentlicht: (2024)
von: Xia, Han, et al.
Veröffentlicht: (2024)
HSKBenchmark: Modeling and Benchmarking Chinese Second Language Acquisition in Large Language Models through Curriculum Tuning
von: Yang, Qihao, et al.
Veröffentlicht: (2025)
von: Yang, Qihao, et al.
Veröffentlicht: (2025)
MedGo: A Chinese Medical Large Language Model
von: Zhang, Haitao, et al.
Veröffentlicht: (2024)
von: Zhang, Haitao, et al.
Veröffentlicht: (2024)
Full-text Error Correction for Chinese Speech Recognition with Large Language Model
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
von: Tang, Zhiyuan, et al.
Veröffentlicht: (2024)
Pillars of Grammatical Error Correction: Comprehensive Inspection Of Contemporary Approaches In The Era of Large Language Models
von: Omelianchuk, Kostiantyn, et al.
Veröffentlicht: (2024)
von: Omelianchuk, Kostiantyn, et al.
Veröffentlicht: (2024)
The Model Agreed, But Didn't Learn: Diagnosing Surface Compliance in Large Language Models
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models
von: Zhang, Che, et al.
Veröffentlicht: (2024)
von: Zhang, Che, et al.
Veröffentlicht: (2024)
Corrective In-Context Learning: Evaluating Self-Correction in Large Language Models
von: Sanz-Guerrero, Mario, et al.
Veröffentlicht: (2025)
von: Sanz-Guerrero, Mario, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalization
von: Liu, Guangliang, et al.
Veröffentlicht: (2025) -
Pragmatic Inference for Moral Reasoning Acquisition: Generalization via Metapragmatic Links
von: Liu, Guangliang, et al.
Veröffentlicht: (2025) -
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
von: Liu, Guangliang, et al.
Veröffentlicht: (2025) -
Discourse Heuristics For Paradoxically Moral Self-Correction
von: Liu, Guangliang, et al.
Veröffentlicht: (2025) -
On the Convergence of Moral Self-Correction in Large Language Models
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)