Self-correction is Not An Innate Capability in Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Guangliang, Qi, Zimo, Zhang, Xitong, Cheng, Lu, Johnson, Kristen Marie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Discourse Heuristics For Paradoxically Moral Self-Correction
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalization
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Smaller Large Language Models Can Do Moral Self-Correction
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models
von: Chen, Bocheng, et al.
Veröffentlicht: (2026)
von: Chen, Bocheng, et al.
Veröffentlicht: (2026)
On the Convergence of Moral Self-Correction in Large Language Models
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Intrinsic Self-correction for Enhanced Morality: An Analysis of Internal Mechanisms and the Superficial Hypothesis
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
On the Intrinsic Self-Correction Capability of LLMs: Uncertainty and Latent Concept
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
Diagnosing the Performance Trade-off in Moral Alignment: A Case Study on Gender Stereotypes
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Pragmatic Inference for Moral Reasoning Acquisition: Generalization via Metapragmatic Links
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)
Towards Understanding Task-agnostic Debiasing Through the Lenses of Intrinsic Bias and Forgetfulness
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
von: Liu, Guangliang, et al.
Veröffentlicht: (2024)
Context-Aware Counterfactual Data Augmentation for Gender Bias Mitigation in Language Models
von: Parihar, Shweta, et al.
Veröffentlicht: (2026)
von: Parihar, Shweta, et al.
Veröffentlicht: (2026)
A Survey to Recent Progress Towards Understanding In-Context Learning
von: Mao, Haitao, et al.
Veröffentlicht: (2024)
von: Mao, Haitao, et al.
Veröffentlicht: (2024)
Refining Salience-Aware Sparse Fine-Tuning Strategies for Language Models
von: Liu, Xinxin, et al.
Veröffentlicht: (2024)
von: Liu, Xinxin, et al.
Veröffentlicht: (2024)
Self-Debias: Self-correcting for Debiasing Large Language Models
von: Feng, Xuan, et al.
Veröffentlicht: (2026)
von: Feng, Xuan, et al.
Veröffentlicht: (2026)
Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
von: Zhang, Yuanchi, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanchi, et al.
Veröffentlicht: (2024)
Small Language Model Can Self-correct
von: Han, Haixia, et al.
Veröffentlicht: (2024)
von: Han, Haixia, et al.
Veröffentlicht: (2024)
ORPP: Self-Optimizing Role-playing Prompts to Enhance Language Model Capabilities
von: Duan, Yifan, et al.
Veröffentlicht: (2025)
von: Duan, Yifan, et al.
Veröffentlicht: (2025)
Hallucination is Inevitable: An Innate Limitation of Large Language Models
von: Xu, Ziwei, et al.
Veröffentlicht: (2024)
von: Xu, Ziwei, et al.
Veröffentlicht: (2024)
Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models
von: Liu, Weize, et al.
Veröffentlicht: (2023)
von: Liu, Weize, et al.
Veröffentlicht: (2023)
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models
von: Song, Yuda, et al.
Veröffentlicht: (2024)
von: Song, Yuda, et al.
Veröffentlicht: (2024)
Confidence Matters: Revisiting Intrinsic Self-Correction Capabilities of Large Language Models
von: Li, Loka, et al.
Veröffentlicht: (2024)
von: Li, Loka, et al.
Veröffentlicht: (2024)
MHPP: Exploring the Capabilities and Limitations of Language Models Beyond Basic Code Generation
von: Dai, Jianbo, et al.
Veröffentlicht: (2024)
von: Dai, Jianbo, et al.
Veröffentlicht: (2024)
Can Large Language Models Handle Discourse Particles? A Case Study of Colloquial Malay
von: Yusoff, Mariah Al Giptiah Binte, et al.
Veröffentlicht: (2026)
von: Yusoff, Mariah Al Giptiah Binte, et al.
Veröffentlicht: (2026)
CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
von: Huang, Shiting, et al.
Veröffentlicht: (2025)
von: Huang, Shiting, et al.
Veröffentlicht: (2025)
Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
von: Huang, Liangjie, et al.
Veröffentlicht: (2025)
von: Huang, Liangjie, et al.
Veröffentlicht: (2025)
How Does Sequence Modeling Architecture Influence Base Capabilities of Pre-trained Language Models? Exploring Key Architecture Design Principles to Avoid Base Capabilities Degradation
von: Lu, Xin, et al.
Veröffentlicht: (2025)
von: Lu, Xin, et al.
Veröffentlicht: (2025)
S$^2$R: Teaching LLMs to Self-verify and Self-correct via Reinforcement Learning
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
von: Ma, Ruotian, et al.
Veröffentlicht: (2025)
Enhancing the Capability and Robustness of Large Language Models through Reinforcement Learning-Driven Query Refinement
von: Wang, Xiaohua, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohua, et al.
Veröffentlicht: (2024)
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
von: Gu, Yiyang, et al.
Veröffentlicht: (2026)
von: Gu, Yiyang, et al.
Veröffentlicht: (2026)
Self-Improving for Zero-Shot Named Entity Recognition with Large Language Models
von: Xie, Tingyu, et al.
Veröffentlicht: (2023)
von: Xie, Tingyu, et al.
Veröffentlicht: (2023)
FALCON: Fine-grained Activation Manipulation by Contrastive Orthogonal Unalignment for Large Language Model
von: Hu, Jinwei, et al.
Veröffentlicht: (2025)
von: Hu, Jinwei, et al.
Veröffentlicht: (2025)
Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
von: Chen, Zehui, et al.
Veröffentlicht: (2023)
Mixture-of-Agents Enhances Large Language Model Capabilities
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
A Survey on Multi-Turn Interaction Capabilities of Large Language Models
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
von: Zhang, Chen, et al.
Veröffentlicht: (2025)
Disinformation Capabilities of Large Language Models
von: Vykopal, Ivan, et al.
Veröffentlicht: (2023)
von: Vykopal, Ivan, et al.
Veröffentlicht: (2023)
Evaluating the Generation Capabilities of Large Chinese Language Models
von: Zeng, Hui, et al.
Veröffentlicht: (2023)
von: Zeng, Hui, et al.
Veröffentlicht: (2023)
Evaluating Accounting Reasoning Capabilities of Large Language Models
von: Zhou, Jie, et al.
Veröffentlicht: (2026)
von: Zhou, Jie, et al.
Veröffentlicht: (2026)
What Are the Odds? Language Models Are Capable of Probabilistic Reasoning
von: Paruchuri, Akshay, et al.
Veröffentlicht: (2024)
von: Paruchuri, Akshay, et al.
Veröffentlicht: (2024)
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
von: Tu, Ruibo, et al.
Veröffentlicht: (2024)
von: Tu, Ruibo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Discourse Heuristics For Paradoxically Moral Self-Correction
von: Liu, Guangliang, et al.
Veröffentlicht: (2025) -
Diagnosing Moral Reasoning Acquisition in Language Models: Pragmatics and Generalization
von: Liu, Guangliang, et al.
Veröffentlicht: (2025) -
Smaller Large Language Models Can Do Moral Self-Correction
von: Liu, Guangliang, et al.
Veröffentlicht: (2024) -
Learning to Diagnose and Correct Errors: Towards Moral Sensitivity Acquisition in Large Language Models
von: Chen, Bocheng, et al.
Veröffentlicht: (2026) -
On the Convergence of Moral Self-Correction in Large Language Models
von: Liu, Guangliang, et al.
Veröffentlicht: (2025)