Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
Fuente:
arXiv
Salvato in:
| Autori principali: | Tutek, Martin, Chaleshtori, Fateme Hashemi, Marasović, Ana, Belinkov, Yonatan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
di: Woo, Jesse, et al.
Pubblicazione: (2025)
di: Woo, Jesse, et al.
Pubblicazione: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
di: Ashuach, Tomer, et al.
Pubblicazione: (2024)
di: Ashuach, Tomer, et al.
Pubblicazione: (2024)
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
di: Chaleshtori, Fateme Hashemi, et al.
Pubblicazione: (2024)
di: Chaleshtori, Fateme Hashemi, et al.
Pubblicazione: (2024)
Teaching People LLM's Errors and Getting it Right
di: Stringham, Nathan, et al.
Pubblicazione: (2025)
di: Stringham, Nathan, et al.
Pubblicazione: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
Reasoning Models Know What's Important, and Encode It in Their Activations
di: Nikankin, Yaniv, et al.
Pubblicazione: (2026)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2026)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026)
di: Gur-Arieh, Yoav, et al.
Pubblicazione: (2026)
Chain-of-Thought Unfaithfulness as Disguised Accuracy
di: Bentham, Oliver, et al.
Pubblicazione: (2024)
di: Bentham, Oliver, et al.
Pubblicazione: (2024)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
di: Hanna, Michael, et al.
Pubblicazione: (2024)
di: Hanna, Michael, et al.
Pubblicazione: (2024)
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
di: Paul, Debjit, et al.
Pubblicazione: (2024)
di: Paul, Debjit, et al.
Pubblicazione: (2024)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2026)
di: Simhi, Adi, et al.
Pubblicazione: (2026)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
Faithful Logical Reasoning via Symbolic Chain-of-Thought
di: Xu, Jundong, et al.
Pubblicazione: (2024)
di: Xu, Jundong, et al.
Pubblicazione: (2024)
Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning
di: Ye, Donald, et al.
Pubblicazione: (2026)
di: Ye, Donald, et al.
Pubblicazione: (2026)
Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
di: Shao, Shun, et al.
Pubblicazione: (2026)
di: Shao, Shun, et al.
Pubblicazione: (2026)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
di: Jia, Jinghan, et al.
Pubblicazione: (2026)
di: Jia, Jinghan, et al.
Pubblicazione: (2026)
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
di: Tanneru, Sree Harsha, et al.
Pubblicazione: (2024)
di: Tanneru, Sree Harsha, et al.
Pubblicazione: (2024)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?
di: Young, Richard J.
Pubblicazione: (2026)
di: Young, Richard J.
Pubblicazione: (2026)
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
di: Meek, Austin, et al.
Pubblicazione: (2025)
di: Meek, Austin, et al.
Pubblicazione: (2025)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
di: Mittal, Avni, et al.
Pubblicazione: (2026)
di: Mittal, Avni, et al.
Pubblicazione: (2026)
FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
di: Gui, Runquan, et al.
Pubblicazione: (2026)
di: Gui, Runquan, et al.
Pubblicazione: (2026)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
di: Hase, Peter, et al.
Pubblicazione: (2026)
di: Hase, Peter, et al.
Pubblicazione: (2026)
A Comprehensive Evaluation of Multilingual Chain-of-Thought Reasoning: Performance, Consistency, and Faithfulness Across Languages
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
di: Zhao, Raoyuan, et al.
Pubblicazione: (2025)
Bounds of Chain-of-Thought Robustness: Reasoning Steps, Embed Norms, and Beyond
di: Wang, Dingzirui, et al.
Pubblicazione: (2025)
di: Wang, Dingzirui, et al.
Pubblicazione: (2025)
A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
di: Lv, Weijiang, et al.
Pubblicazione: (2026)
di: Lv, Weijiang, et al.
Pubblicazione: (2026)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
di: Carmeli, Boaz, et al.
Pubblicazione: (2024)
di: Carmeli, Boaz, et al.
Pubblicazione: (2024)
Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
di: Young, Richard J.
Pubblicazione: (2026)
di: Young, Richard J.
Pubblicazione: (2026)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
di: Zaman, Kerem, et al.
Pubblicazione: (2025)
di: Zaman, Kerem, et al.
Pubblicazione: (2025)
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
Are formal and functional linguistic mechanisms dissociated in language models?
di: Hanna, Michael, et al.
Pubblicazione: (2025)
di: Hanna, Michael, et al.
Pubblicazione: (2025)
Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
SAEs Are Good for Steering -- If You Select the Right Features
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
Documenti analoghi
-
BriefMe: A Legal NLP Benchmark for Assisting with Legal Briefs
di: Woo, Jesse, et al.
Pubblicazione: (2025) -
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
di: Ashuach, Tomer, et al.
Pubblicazione: (2024) -
On Evaluating Explanation Utility for Human-AI Decision Making in NLP
di: Chaleshtori, Fateme Hashemi, et al.
Pubblicazione: (2024) -
Teaching People LLM's Errors and Getting it Right
di: Stringham, Nathan, et al.
Pubblicazione: (2025) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)