Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures
Fuente:
arXiv
Saved in:
| Main Authors: | Somov, Oleg, Chaichuk, Mikhail, Seleznyov, Mikhail, Panchenko, Alexander, Tutubalina, Elena |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
by: Seleznyov, Mikhail, et al.
Published: (2025)
by: Seleznyov, Mikhail, et al.
Published: (2025)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
by: Seleznyov, Mikhail, et al.
Published: (2026)
by: Seleznyov, Mikhail, et al.
Published: (2026)
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
by: Alekseev, Artem, et al.
Published: (2025)
by: Alekseev, Artem, et al.
Published: (2025)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
by: Vazhentsev, Artem, et al.
Published: (2026)
by: Vazhentsev, Artem, et al.
Published: (2026)
Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models
by: Chaichuk, Mikhail, et al.
Published: (2025)
by: Chaichuk, Mikhail, et al.
Published: (2025)
Confidence Estimation for Error Detection in Text-to-SQL Systems
by: Somov, Oleg, et al.
Published: (2025)
by: Somov, Oleg, et al.
Published: (2025)
RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation
by: Bondarenko, Ivan, et al.
Published: (2026)
by: Bondarenko, Ivan, et al.
Published: (2026)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
by: Mittal, Avni, et al.
Published: (2026)
by: Mittal, Avni, et al.
Published: (2026)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
by: Braslavski, Pavel, et al.
Published: (2026)
by: Braslavski, Pavel, et al.
Published: (2026)
Project Ariadne: A Structural Causal Framework for Auditing Faithfulness in LLM Agents
by: Khanzadeh, Sourena
Published: (2026)
by: Khanzadeh, Sourena
Published: (2026)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
by: Korznikov, Anton, et al.
Published: (2025)
by: Korznikov, Anton, et al.
Published: (2025)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
by: Swaroop, Anand, et al.
Published: (2025)
by: Swaroop, Anand, et al.
Published: (2025)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
by: Larionov, Daniil, et al.
Published: (2024)
by: Larionov, Daniil, et al.
Published: (2024)
Hybrid LLM/Rule-based Approaches to Business Insights Generation from Structured Data
by: Vertsel, Aliaksei, et al.
Published: (2024)
by: Vertsel, Aliaksei, et al.
Published: (2024)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
by: Shen, Xu, et al.
Published: (2025)
by: Shen, Xu, et al.
Published: (2025)
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
by: Sakhovskiy, Andrey, et al.
Published: (2025)
by: Sakhovskiy, Andrey, et al.
Published: (2025)
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
by: Lv, Weijiang, et al.
Published: (2026)
by: Lv, Weijiang, et al.
Published: (2026)
Sparse Prefix Caching for Hybrid and Recurrent LLM Serving
by: Shirokikh, Mikhail, et al.
Published: (2026)
by: Shirokikh, Mikhail, et al.
Published: (2026)
S3: A Simple Strong Sample-effective Multimodal Dialog System
by: Rykov, Elisei, et al.
Published: (2024)
by: Rykov, Elisei, et al.
Published: (2024)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
by: Afonin, Nikita, et al.
Published: (2025)
by: Afonin, Nikita, et al.
Published: (2025)
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
by: Dementieva, Daryna, et al.
Published: (2024)
by: Dementieva, Daryna, et al.
Published: (2024)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
by: Pletenev, Sergey, et al.
Published: (2025)
by: Pletenev, Sergey, et al.
Published: (2025)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
by: Murphy, Alexander, et al.
Published: (2025)
by: Murphy, Alexander, et al.
Published: (2025)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
by: Hase, Peter, et al.
Published: (2026)
by: Hase, Peter, et al.
Published: (2026)
Framework GNN-AID: Graph Neural Network Analysis Interpretation and Defense
by: Lukyanov, Kirill, et al.
Published: (2025)
by: Lukyanov, Kirill, et al.
Published: (2025)
Orientability of Causal Relations in Time Series using Summary Causal Graphs and Faithful Distributions
by: Loranchet, Timothée, et al.
Published: (2025)
by: Loranchet, Timothée, et al.
Published: (2025)
Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness
by: Li, Jiachun, et al.
Published: (2024)
by: Li, Jiachun, et al.
Published: (2024)
Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
by: Yang, Chenxu, et al.
Published: (2025)
by: Yang, Chenxu, et al.
Published: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
Probabilistic Verification of Voice Anti-Spoofing Models
by: Kushnir, Evgeny, et al.
Published: (2026)
by: Kushnir, Evgeny, et al.
Published: (2026)
LLM-Guided Prompt Evolution for Password Guessing
by: Mazin, Vladimir A., et al.
Published: (2026)
by: Mazin, Vladimir A., et al.
Published: (2026)
Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
by: Young, Richard J.
Published: (2026)
by: Young, Richard J.
Published: (2026)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
by: Arcuschin, Iván, et al.
Published: (2025)
by: Arcuschin, Iván, et al.
Published: (2025)
Equation discovery framework EPDE: Towards a better equation discovery
by: Maslyaev, Mikhail, et al.
Published: (2024)
by: Maslyaev, Mikhail, et al.
Published: (2024)
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
Intermediate Languages Matter: Formal Choice Drives Neurosymbolic LLM Reasoning
by: Beiser, Alexander, et al.
Published: (2025)
by: Beiser, Alexander, et al.
Published: (2025)
GameUIAgent: An LLM-Powered Framework for Automated Game UI Design with Structured Intermediate Representation
by: Zeng, Wei, et al.
Published: (2026)
by: Zeng, Wei, et al.
Published: (2026)
When Bits Break Recourse: Counterfactual-Faithful Quantization
by: Yahyati, Chaymae, et al.
Published: (2026)
by: Yahyati, Chaymae, et al.
Published: (2026)
ReCo: Reliable Causal Chain Reasoning via Structural Causal Recurrent Neural Networks
by: Xiong, Kai, et al.
Published: (2022)
by: Xiong, Kai, et al.
Published: (2022)
Similar Items
-
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
by: Seleznyov, Mikhail, et al.
Published: (2025) -
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
by: Seleznyov, Mikhail, et al.
Published: (2026) -
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
by: Alekseev, Artem, et al.
Published: (2025) -
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026) -
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
by: Vazhentsev, Artem, et al.
Published: (2026)