Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Arcuschin, Iván, Janiak, Jett, Krzyzanowski, Robert, Rajamanoharan, Senthooran, Nanda, Neel, Conmy, Arthur |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
von: Macar, Uzay, et al.
Veröffentlicht: (2025)
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
von: Ferrando, Javier, et al.
Veröffentlicht: (2024)
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025)
Understanding Reasoning in Thinking Language Models via Steering Vectors
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Base Models Know How to Reason, Thinking Models Learn When
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
von: Meek, Austin, et al.
Veröffentlicht: (2025)
von: Meek, Austin, et al.
Veröffentlicht: (2025)
Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
von: Lieberum, Tom, et al.
Veröffentlicht: (2024)
Steering Out-of-Distribution Generalization with Concept Ablation Fine-Tuning
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
von: Casademunt, Helena, et al.
Veröffentlicht: (2025)
Emergent Misalignment is Easy, Narrow Misalignment is Hard
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
von: Soligo, Anna, et al.
Veröffentlicht: (2026)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Improving Dictionary Learning with Gated Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Scaling sparse feature circuit finding for in-context learning
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
von: Kharlapenko, Dmitrii, et al.
Veröffentlicht: (2025)
Subliminal Learning Is Steering Vector Distillation
von: Blank, Camila, et al.
Veröffentlicht: (2026)
von: Blank, Camila, et al.
Veröffentlicht: (2026)
Simple Mechanistic Explanations for Out-Of-Context Reasoning
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
von: Wang, Atticus, et al.
Veröffentlicht: (2025)
Automatically Finding Reward Model Biases
von: Wang, Atticus, et al.
Veröffentlicht: (2026)
von: Wang, Atticus, et al.
Veröffentlicht: (2026)
Model Organisms for Emergent Misalignment
von: Turner, Edward, et al.
Veröffentlicht: (2025)
von: Turner, Edward, et al.
Veröffentlicht: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Building Production-Ready Probes For Gemini
von: Kramár, János, et al.
Veröffentlicht: (2026)
von: Kramár, János, et al.
Veröffentlicht: (2026)
How Well Do Models Follow Their Constitutions?
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
von: Jakkli, Arya, et al.
Veröffentlicht: (2026)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
Dense SAE Latents Are Features, Not Bugs
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Sun, Xiaoqing, et al.
Veröffentlicht: (2025)
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
von: Zhang, Fred, et al.
Veröffentlicht: (2023)
Improving Steering Vectors by Targeting Sparse Autoencoder Features
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
von: Chalnev, Sviatoslav, et al.
Veröffentlicht: (2024)
Fractured Chain-of-Thought Reasoning
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
von: Liao, Baohao, et al.
Veröffentlicht: (2025)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
von: Sun, Jingyi, et al.
Veröffentlicht: (2026)
von: Sun, Jingyi, et al.
Veröffentlicht: (2026)
Eliciting Secret Knowledge from Language Models
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
von: Cywiński, Bartosz, et al.
Veröffentlicht: (2025)
Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
von: Rajamanoharan, Senthooran, et al.
Veröffentlicht: (2024)
Diffusion of Thoughts: Chain-of-Thought Reasoning in Diffusion Language Models
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
von: Ye, Jiacheng, et al.
Veröffentlicht: (2024)
When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
von: Emmons, Scott, et al.
Veröffentlicht: (2025)
von: Emmons, Scott, et al.
Veröffentlicht: (2025)
Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
von: Young, Richard J.
Veröffentlicht: (2026)
von: Young, Richard J.
Veröffentlicht: (2026)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
von: Xie, Zhuohan, et al.
Veröffentlicht: (2025)
von: Xie, Zhuohan, et al.
Veröffentlicht: (2025)
Compositional Reasoning with Transformers, RNNs, and Chain of Thought
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
von: Yehudai, Gilad, et al.
Veröffentlicht: (2025)
Scalable Chain of Thoughts via Elastic Reasoning
von: Xu, Yuhui, et al.
Veröffentlicht: (2025)
von: Xu, Yuhui, et al.
Veröffentlicht: (2025)
Long Chain-of-Thought Reasoning Across Languages
von: Barua, Josh, et al.
Veröffentlicht: (2025)
von: Barua, Josh, et al.
Veröffentlicht: (2025)
Interpreting Attention Layer Outputs with Sparse Autoencoders
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
von: Kissane, Connor, et al.
Veröffentlicht: (2024)
Value-Guided Search for Efficient Chain-of-Thought Reasoning
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
von: Wang, Kaiwen, et al.
Veröffentlicht: (2025)
Enhancing Generalization in Chain of Thought Reasoning for Smaller Models
von: Yin, Maxwell J., et al.
Veröffentlicht: (2025)
von: Yin, Maxwell J., et al.
Veröffentlicht: (2025)
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
von: Hu, Lijie, et al.
Veröffentlicht: (2024)
von: Hu, Lijie, et al.
Veröffentlicht: (2024)
Stepwise Penalization for Length-Efficient Chain-of-Thought Reasoning
von: Li, Xintong, et al.
Veröffentlicht: (2026)
von: Li, Xintong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Thought Branches: Interpreting LLM Reasoning Requires Resampling
von: Macar, Uzay, et al.
Veröffentlicht: (2025) -
Do I Know This Entity? Knowledge Awareness and Hallucinations in Language Models
von: Ferrando, Javier, et al.
Veröffentlicht: (2024) -
Thought Anchors: Which LLM Reasoning Steps Matter?
von: Bogdan, Paul C., et al.
Veröffentlicht: (2025) -
Understanding Reasoning in Thinking Language Models via Steering Vectors
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025) -
Base Models Know How to Reason, Thinking Models Learn When
von: Venhoff, Constantin, et al.
Veröffentlicht: (2025)