FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
Fuente:
arXiv
Saved in:
| Main Authors: | Swaroop, Anand, Nallani, Akshat, Uboweja, Saksham, Uzdenova, Adiliia, Nguyen, Michael, Zhu, Kevin, Dev, Sunishchal, Panda, Ashwinee, Sharma, Vasu, Chaudhary, Maheep |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025)
by: Batra, Shourya, et al.
Published: (2025)
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
by: Egbuna, Nathan, et al.
Published: (2025)
by: Egbuna, Nathan, et al.
Published: (2025)
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
by: Chaturvedi, Isha, et al.
Published: (2025)
by: Chaturvedi, Isha, et al.
Published: (2025)
Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Broken Chains: The Cost of Incomplete Reasoning in LLMs
by: Su, Ian, et al.
Published: (2026)
by: Su, Ian, et al.
Published: (2026)
Alignment-Constrained Dynamic Pruning for LLMs: Identifying and Preserving Alignment-Critical Circuits
by: Patel, Dev, et al.
Published: (2025)
by: Patel, Dev, et al.
Published: (2025)
PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases
by: Vuddanti, Sri Vatsa, et al.
Published: (2025)
by: Vuddanti, Sri Vatsa, et al.
Published: (2025)
Thought Experiments in Design Fiction for Visualization
by: Panda, Swaroop
Published: (2024)
by: Panda, Swaroop
Published: (2024)
Modeling and Predicting Multi-Turn Answer Instability in Large Language Models
by: He, Jiahang, et al.
Published: (2025)
by: He, Jiahang, et al.
Published: (2025)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
by: Chaudhary, Siddharth, et al.
Published: (2025)
by: Chaudhary, Siddharth, et al.
Published: (2025)
Inference-Time Chain-of-Thought Pruning with Latent Informativeness Signals
by: Li, Sophie, et al.
Published: (2025)
by: Li, Sophie, et al.
Published: (2025)
In-Context Environments Induce Evaluation-Awareness in Language Models
by: Chaudhary, Maheep
Published: (2026)
by: Chaudhary, Maheep
Published: (2026)
Optimizing Chain-of-Thought Confidence via Topological and Dirichlet Risk Analysis
by: More, Abhishek, et al.
Published: (2025)
by: More, Abhishek, et al.
Published: (2025)
DuoLens: A Framework for Robust Detection of Machine-Generated Multilingual Text and Code
by: Agrawal, Shriyansh, et al.
Published: (2025)
by: Agrawal, Shriyansh, et al.
Published: (2025)
ProMoral-Bench: Evaluating Prompting Strategies for Moral Reasoning and Safety in LLMs
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
by: Thomas, Rohan Subramanian, et al.
Published: (2026)
A Few Bad Neurons: Isolating and Surgically Correcting Sycophancy
by: O'Brien, Claire, et al.
Published: (2026)
by: O'Brien, Claire, et al.
Published: (2026)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
by: Chaudhary, Maheep, et al.
Published: (2024)
by: Chaudhary, Maheep, et al.
Published: (2024)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
Mechanistic origins of catastrophic forgetting: why RL preserves circuits better than SFT?
by: Nunez, Jeanmely Rojas, et al.
Published: (2026)
by: Nunez, Jeanmely Rojas, et al.
Published: (2026)
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
by: Rana, Manik, et al.
Published: (2025)
by: Rana, Manik, et al.
Published: (2025)
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
by: Afonin, Nikita, et al.
Published: (2025)
by: Afonin, Nikita, et al.
Published: (2025)
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
by: Paul, Debjit, et al.
Published: (2024)
by: Paul, Debjit, et al.
Published: (2024)
Confidential FRIT via Homomorphic Encryption
by: Hoshino, Haruki, et al.
Published: (2025)
by: Hoshino, Haruki, et al.
Published: (2025)
LLMs' ways of seeing User Personas
by: Panda, Swaroop
Published: (2024)
by: Panda, Swaroop
Published: (2024)
Making AI Agents Evaluate Misleading Charts without Nudging
by: Panda, Swaroop
Published: (2026)
by: Panda, Swaroop
Published: (2026)
Look into your Heart -- Prototypes for a Speculative Design Exploration of Personal Heart Rate Visualization
by: Panda, Swaroop
Published: (2025)
by: Panda, Swaroop
Published: (2025)
A Framework for LLM-powered Design Assistants
by: Panda, Swaroop
Published: (2025)
by: Panda, Swaroop
Published: (2025)
Physical therapeutic elements in stage medical recovery of puerperas
by: Zukhra Kh. Uzdenova
Published: (2021)
by: Zukhra Kh. Uzdenova
Published: (2021)
Physical therapeutic factors in stage medical rehabilitation of puerperas with perineal wounds after fetal vacuum extraction
by: Zukhra Kh. Uzdenova
Published: (2021)
by: Zukhra Kh. Uzdenova
Published: (2021)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
by: Jia, Jinghan, et al.
Published: (2026)
by: Jia, Jinghan, et al.
Published: (2026)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
by: Hase, Peter, et al.
Published: (2026)
by: Hase, Peter, et al.
Published: (2026)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
by: Arcuschin, Iván, et al.
Published: (2025)
by: Arcuschin, Iván, et al.
Published: (2025)
@GrokSet: multi-party Human-LLM Interactions in Social Media
by: Migliarini, Matteo, et al.
Published: (2026)
by: Migliarini, Matteo, et al.
Published: (2026)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
Sumudu Neural Operator for ODEs and PDEs
by: Zelenskiy, Ben, et al.
Published: (2025)
by: Zelenskiy, Ben, et al.
Published: (2025)
Causal Sufficiency and Necessity Improves Chain-of-Thought Reasoning
by: Yu, Xiangning, et al.
Published: (2025)
by: Yu, Xiangning, et al.
Published: (2025)
Freespace twistronics for optical supertopologies
by: Dev, Vasu, et al.
Published: (2025)
by: Dev, Vasu, et al.
Published: (2025)
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
Similar Items
-
SALT: Steering Activations towards Leakage-free Thinking in Chain of Thought
by: Batra, Shourya, et al.
Published: (2025) -
Amortized Latent Steering: Low-Cost Alternative to Test-Time Optimization
by: Egbuna, Nathan, et al.
Published: (2025) -
Peek-a-Boo Reasoning: Contrastive Region Masking in MLLMs
by: Chaturvedi, Isha, et al.
Published: (2025) -
Evaluation Awareness Scales Predictably in Open-Weights Large Language Models
by: Chaudhary, Maheep, et al.
Published: (2025) -
Broken Chains: The Cost of Incomplete Reasoning in LLMs
by: Su, Ian, et al.
Published: (2026)