Do Models Know Why They Changed Their Mind? Interpretability and Faithfulness of Chain-of-Thought Under Knowledge Conflict
Fuente:
arXiv
Saved in:
| Main Author: | Venkata, Pruthvinath Jeripity |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Three Regimes of Context-Parametric Conflict: A Predictive Framework and Empirical Validation
by: Venkata, Pruthvinath Jeripity
Published: (2026)
by: Venkata, Pruthvinath Jeripity
Published: (2026)
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025)
by: Balasubramanian, Sriram, et al.
Published: (2025)
The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies
by: Garcia, Gabriel
Published: (2026)
by: Garcia, Gabriel
Published: (2026)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)
by: Oketunji, Abiodun Finbarrs
Published: (2023)
Applied Explainability for Large Language Models: A Comparative Study
by: Kancharla, Venkata Abhinandan
Published: (2026)
by: Kancharla, Venkata Abhinandan
Published: (2026)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
by: Höth, Max Henning, et al.
Published: (2026)
by: Höth, Max Henning, et al.
Published: (2026)
Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency
by: Hong, Colin, et al.
Published: (2025)
by: Hong, Colin, et al.
Published: (2025)
Do Biased Models Have Biased Thoughts?
by: Rajwal, Swati, et al.
Published: (2025)
by: Rajwal, Swati, et al.
Published: (2025)
The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level
by: Herbst, Jeremy, et al.
Published: (2026)
by: Herbst, Jeremy, et al.
Published: (2026)
NLP Case Study on Predicting the Before and After of the Ukraine-Russia and Hamas-Israel Conflicts
by: Miner, Jordan, et al.
Published: (2024)
by: Miner, Jordan, et al.
Published: (2024)
Do Personality Traits Interfere? Geometric Limitations of Steering in Large Language Models
by: Bhandari, Pranav, et al.
Published: (2026)
by: Bhandari, Pranav, et al.
Published: (2026)
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
by: Gao, Yutong, et al.
Published: (2026)
by: Gao, Yutong, et al.
Published: (2026)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
by: Yordanov, Yordan, et al.
Published: (2026)
by: Yordanov, Yordan, et al.
Published: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
by: Oketunji, Abiodun Finbarrs, et al.
Published: (2023)
Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
by: Nainani, Jatin, et al.
Published: (2024)
by: Nainani, Jatin, et al.
Published: (2024)
Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
by: Labadie-Tamayo, Roberto, et al.
Published: (2025)
by: Labadie-Tamayo, Roberto, et al.
Published: (2025)
DynaSemble: Dynamic Ensembling of Textual and Structure-Based Models for Knowledge Graph Completion
by: Nandi, Ananjan, et al.
Published: (2023)
by: Nandi, Ananjan, et al.
Published: (2023)
Thought-Like-Pro: Enhancing Reasoning of Large Language Models through Self-Driven Prolog-based Chain-of-Thought
by: Tan, Xiaoyu, et al.
Published: (2024)
by: Tan, Xiaoyu, et al.
Published: (2024)
RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models
by: Yan, Qihang, et al.
Published: (2025)
by: Yan, Qihang, et al.
Published: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
by: Fadli, Samih
Published: (2025)
by: Fadli, Samih
Published: (2025)
Knowledge Graph Embeddings: A Comprehensive Survey on Capturing Relation Properties
by: Niu, Guanglin
Published: (2024)
by: Niu, Guanglin
Published: (2024)
On Measuring Faithfulness or Self-consistency of Natural Language Explanations
by: Parcalabescu, Letitia, et al.
Published: (2023)
by: Parcalabescu, Letitia, et al.
Published: (2023)
CodingTeachLLM: Empowering LLM's Coding Ability via AST Prior Knowledge
by: Chen, Zhangquan, et al.
Published: (2024)
by: Chen, Zhangquan, et al.
Published: (2024)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
by: Anderson, Samuel Cyrenius
Published: (2026)
by: Anderson, Samuel Cyrenius
Published: (2026)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
by: Cho, Seonglae, et al.
Published: (2026)
by: Cho, Seonglae, et al.
Published: (2026)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
by: Zhou, Tianyang, et al.
Published: (2026)
by: Zhou, Tianyang, et al.
Published: (2026)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
by: Hanna, Michael, et al.
Published: (2024)
by: Hanna, Michael, et al.
Published: (2024)
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
by: Cacioli, Jon-Paul
Published: (2026)
by: Cacioli, Jon-Paul
Published: (2026)
Change Is the Only Constant: Dynamic LLM Slicing based on Layer Redundancy
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
by: Dumitru, Razvan-Gabriel, et al.
Published: (2024)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
by: Lan, Guangchen, et al.
Published: (2025)
by: Lan, Guangchen, et al.
Published: (2025)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
by: Yu, Zony, et al.
Published: (2025)
by: Yu, Zony, et al.
Published: (2025)
Large Language Models Can Better Understand Knowledge Graphs Than We Thought
by: Dai, Xinbang, et al.
Published: (2024)
by: Dai, Xinbang, et al.
Published: (2024)
Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
by: Mircea, Andrei, et al.
Published: (2025)
by: Mircea, Andrei, et al.
Published: (2025)
REAP the Experts: Why Pruning Prevails for One-Shot MoE compression
by: Lasby, Mike, et al.
Published: (2025)
by: Lasby, Mike, et al.
Published: (2025)
Random-Set Large Language Models
by: Mubashar, Muhammad, et al.
Published: (2025)
by: Mubashar, Muhammad, et al.
Published: (2025)
Improving Language Models with Intentional Analysis
by: Yin, Yuwei, et al.
Published: (2025)
by: Yin, Yuwei, et al.
Published: (2025)
Danoliteracy of Generative Large Language Models
by: Holm, Søren Vejlgaard, et al.
Published: (2024)
by: Holm, Søren Vejlgaard, et al.
Published: (2024)
SWI: Speaking with Intent in Large Language Models
by: Yin, Yuwei, et al.
Published: (2025)
by: Yin, Yuwei, et al.
Published: (2025)
Uncovering Biases with Reflective Large Language Models
by: Chang, Edward Y.
Published: (2024)
by: Chang, Edward Y.
Published: (2024)
Similar Items
-
Three Regimes of Context-Parametric Conflict: A Predictive Framework and Empirical Validation
by: Venkata, Pruthvinath Jeripity
Published: (2026) -
Why Models Know But Don't Say: Chain-of-Thought Faithfulness Divergence Between Thinking Tokens and Answers in Open-Weight Reasoning Models
by: Young, Richard J.
Published: (2026) -
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
by: Balasubramanian, Sriram, et al.
Published: (2025) -
The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies
by: Garcia, Gabriel
Published: (2026) -
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
by: Oketunji, Abiodun Finbarrs
Published: (2023)