Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yu-Ting, Chang, Fu-Chieh, Shu, Yu-En, Shih, Hui-Ying, Wu, Pei-Yuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner
by: Chang, Fu-Chieh, et al.
Published: (2024)
by: Chang, Fu-Chieh, et al.
Published: (2024)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
by: Yang, Xin, et al.
Published: (2026)
by: Yang, Xin, et al.
Published: (2026)
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
by: Yu, Zeping, et al.
Published: (2024)
by: Yu, Zeping, et al.
Published: (2024)
Understanding the Dark Side of LLMs' Intrinsic Self-Correction
by: Zhang, Qingjie, et al.
Published: (2024)
by: Zhang, Qingjie, et al.
Published: (2024)
Unraveling Arithmetic in Large Language Models: The Role of Algebraic Structures
by: Chang, Fu-Chieh, et al.
Published: (2024)
by: Chang, Fu-Chieh, et al.
Published: (2024)
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
by: Chang, Kai-Po, et al.
Published: (2025)
by: Chang, Kai-Po, et al.
Published: (2025)
On the Intrinsic Self-Correction Capability of LLMs: Uncertainty and Latent Concept
by: Liu, Guangliang, et al.
Published: (2024)
by: Liu, Guangliang, et al.
Published: (2024)
Unveiling the Latent Directions of Reflection in Large Language Models
by: Chang, Fu-Chieh, et al.
Published: (2025)
by: Chang, Fu-Chieh, et al.
Published: (2025)
Transfer-Prompting: Enhancing Cross-Task Adaptation in Large Language Models via Dual-Stage Prompts Optimization
by: Chang, Yupeng, et al.
Published: (2025)
by: Chang, Yupeng, et al.
Published: (2025)
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
by: Tian, Ye, et al.
Published: (2024)
by: Tian, Ye, et al.
Published: (2024)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
Dissecting Bias in LLMs: A Mechanistic Interpretability Perspective
by: Chandna, Bhavik, et al.
Published: (2025)
by: Chandna, Bhavik, et al.
Published: (2025)
Can LLMs Correct Themselves? A Benchmark of Self-Correction in LLMs
by: Tie, Guiyao, et al.
Published: (2025)
by: Tie, Guiyao, et al.
Published: (2025)
Revise, Don't Freeze: Sampler-Matched Training for Self-Correcting Masked Diffusion Language Models
by: Yu, Longxuan, et al.
Published: (2026)
by: Yu, Longxuan, et al.
Published: (2026)
Mechanistic Interpretability of ASR models using Sparse Autoencoders
by: Pluth, Dan, et al.
Published: (2026)
by: Pluth, Dan, et al.
Published: (2026)
Mechanistic Interpretability of Cognitive Complexity in LLMs via Linear Probing using Bloom's Taxonomy
by: Raimondi, Bianca, et al.
Published: (2026)
by: Raimondi, Bianca, et al.
Published: (2026)
Is Inference Mediated by Distinct Semantic Structures in LLMs? A Mechanistic Interpretation
by: Aljaafari, Nura, et al.
Published: (2026)
by: Aljaafari, Nura, et al.
Published: (2026)
PromptEmbedder:: Efficient and Transferable Text Embedding via Dual-LLM Soft Prompting
by: Tsai, Yu-Che, et al.
Published: (2026)
by: Tsai, Yu-Che, et al.
Published: (2026)
Supervised Optimism Correction: Be Confident When LLMs Are Sure
by: Zhang, Junjie, et al.
Published: (2025)
by: Zhang, Junjie, et al.
Published: (2025)
Analysing Moral Bias in Finetuned LLMs through Mechanistic Interpretability
by: Raimondi, Bianca, et al.
Published: (2025)
by: Raimondi, Bianca, et al.
Published: (2025)
MedReflect: Teaching Medical LLMs to Self-Improve via Reflective Correction
by: Huang, Yue, et al.
Published: (2025)
by: Huang, Yue, et al.
Published: (2025)
Understanding and Mitigating Gender Bias in LLMs via Interpretable Neuron Editing
by: Yu, Zeping, et al.
Published: (2025)
by: Yu, Zeping, et al.
Published: (2025)
Large Language Models have Intrinsic Self-Correction Ability
by: Liu, Dancheng, et al.
Published: (2024)
by: Liu, Dancheng, et al.
Published: (2024)
Towards Generalist Prompting for Large Language Models by Mental Models
by: Guan, Haoxiang, et al.
Published: (2024)
by: Guan, Haoxiang, et al.
Published: (2024)
Improving TCM Question Answering through Tree-Organized Self-Reflective Retrieval with LLMs
by: Liu, Chang, et al.
Published: (2025)
by: Liu, Chang, et al.
Published: (2025)
ReDeEP: Detecting Hallucination in Retrieval-Augmented Generation via Mechanistic Interpretability
by: Sun, Zhongxiang, et al.
Published: (2024)
by: Sun, Zhongxiang, et al.
Published: (2024)
Learning Intrinsic Dimension via Information Bottleneck for Explainable Aspect-based Sentiment Analysis
by: Cheng, Zhenxiao, et al.
Published: (2024)
by: Cheng, Zhenxiao, et al.
Published: (2024)
HyperDAS: Towards Automating Mechanistic Interpretability with Hypernetworks
by: Sun, Jiuding, et al.
Published: (2025)
by: Sun, Jiuding, et al.
Published: (2025)
Toward a Theory of Generalizability in LLM Mechanistic Interpretability Research
by: Trott, Sean
Published: (2025)
by: Trott, Sean
Published: (2025)
From Syntax to Emotion: A Mechanistic Analysis of Emotion Inference in LLMs
by: Shu, Bangzhao, et al.
Published: (2026)
by: Shu, Bangzhao, et al.
Published: (2026)
DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders
by: Wang, Xu, et al.
Published: (2026)
by: Wang, Xu, et al.
Published: (2026)
Mechanistic Interpretability of Large-Scale Counting in LLMs through a System-2 Strategy
by: Hasani, Hosein, et al.
Published: (2026)
by: Hasani, Hosein, et al.
Published: (2026)
Towards Self-Improvement of LLMs via MCTS: Leveraging Stepwise Knowledge with Curriculum Preference Learning
by: Wang, Xiyao, et al.
Published: (2024)
by: Wang, Xiyao, et al.
Published: (2024)
Towards Ethical Multi-Agent Systems of Large Language Models: A Mechanistic Interpretability Perspective
by: Lee, Jae Hee, et al.
Published: (2025)
by: Lee, Jae Hee, et al.
Published: (2025)
An Exploratory Framework for Future SETI Applications: Detecting Generative Reactivity via Language Models
by: Yu, Po-Chieh
Published: (2025)
by: Yu, Po-Chieh
Published: (2025)
Stream: Scaling up Mechanistic Interpretability to Long Context in LLMs via Sparse Attention
by: Rosser, J, et al.
Published: (2025)
by: Rosser, J, et al.
Published: (2025)
Minimal and Mechanistic Conditions for Behavioral Self-Awareness in LLMs
by: Bozoukov, Matthew, et al.
Published: (2025)
by: Bozoukov, Matthew, et al.
Published: (2025)
MinPrompt: Graph-based Minimal Prompt Data Augmentation for Few-shot Question Answering
by: Chen, Xiusi, et al.
Published: (2023)
by: Chen, Xiusi, et al.
Published: (2023)
A Theoretical Framework for OOD Robustness in Transformers using Gevrey Classes
by: Wang, Yu, et al.
Published: (2025)
by: Wang, Yu, et al.
Published: (2025)
Preference Heads in Large Language Models: A Mechanistic Framework for Interpretable Personalization
by: Zhang, Weixu, et al.
Published: (2026)
by: Zhang, Weixu, et al.
Published: (2026)
Similar Items
-
RL-STaR: Theoretical Analysis of Reinforcement Learning Frameworks for Self-Taught Reasoner
by: Chang, Fu-Chieh, et al.
Published: (2024) -
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
by: Yang, Xin, et al.
Published: (2026) -
Understanding Multimodal LLMs: the Mechanistic Interpretability of Llava in Visual Question Answering
by: Yu, Zeping, et al.
Published: (2024) -
Understanding the Dark Side of LLMs' Intrinsic Self-Correction
by: Zhang, Qingjie, et al.
Published: (2024) -
Unraveling Arithmetic in Large Language Models: The Role of Algebraic Structures
by: Chang, Fu-Chieh, et al.
Published: (2024)