Breaking the Chain: A Causal Analysis of LLM Faithfulness to Intermediate Structures
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Somov, Oleg, Chaichuk, Mikhail, Seleznyov, Mikhail, Panchenko, Alexander, Tutubalina, Elena |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2025)
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2025)
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026)
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026)
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
von: Alekseev, Artem, et al.
Veröffentlicht: (2025)
von: Alekseev, Artem, et al.
Veröffentlicht: (2025)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)
Prompt to Polyp: Medical Text-Conditioned Image Synthesis with Diffusion Models
von: Chaichuk, Mikhail, et al.
Veröffentlicht: (2025)
von: Chaichuk, Mikhail, et al.
Veröffentlicht: (2025)
Confidence Estimation for Error Detection in Text-to-SQL Systems
von: Somov, Oleg, et al.
Veröffentlicht: (2025)
von: Somov, Oleg, et al.
Veröffentlicht: (2025)
RaguTeam at SemEval-2026 Task 8: Meno and Friends in a Judge-Orchestrated LLM Ensemble for Faithful Multi-Turn Response Generation
von: Bondarenko, Ivan, et al.
Veröffentlicht: (2026)
von: Bondarenko, Ivan, et al.
Veröffentlicht: (2026)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
von: Mittal, Avni, et al.
Veröffentlicht: (2026)
The Chronicles of RiDiC: Generating Datasets with Controlled Popularity Distribution for Long-form Factuality Evaluation
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
von: Braslavski, Pavel, et al.
Veröffentlicht: (2026)
Project Ariadne: A Structural Causal Framework for Auditing Faithfulness in LLM Agents
von: Khanzadeh, Sourena
Veröffentlicht: (2026)
von: Khanzadeh, Sourena
Veröffentlicht: (2026)
The Rogue Scalpel: Activation Steering Compromises LLM Safety
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
von: Korznikov, Anton, et al.
Veröffentlicht: (2025)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
von: Swaroop, Anand, et al.
Veröffentlicht: (2025)
von: Swaroop, Anand, et al.
Veröffentlicht: (2025)
xCOMET-lite: Bridging the Gap Between Efficiency and Quality in Learned MT Evaluation Metrics
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
von: Larionov, Daniil, et al.
Veröffentlicht: (2024)
Hybrid LLM/Rule-based Approaches to Business Insights Generation from Structured Data
von: Vertsel, Aliaksei, et al.
Veröffentlicht: (2024)
von: Vertsel, Aliaksei, et al.
Veröffentlicht: (2024)
FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
von: Shen, Xu, et al.
Veröffentlicht: (2025)
von: Shen, Xu, et al.
Veröffentlicht: (2025)
BALI: Enhancing Biomedical Language Representations through Knowledge Graph and Language Model Alignment
von: Sakhovskiy, Andrey, et al.
Veröffentlicht: (2025)
von: Sakhovskiy, Andrey, et al.
Veröffentlicht: (2025)
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
von: Lv, Weijiang, et al.
Veröffentlicht: (2026)
Sparse Prefix Caching for Hybrid and Recurrent LLM Serving
von: Shirokikh, Mikhail, et al.
Veröffentlicht: (2026)
von: Shirokikh, Mikhail, et al.
Veröffentlicht: (2026)
S3: A Simple Strong Sample-effective Multimodal Dialog System
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
von: Rykov, Elisei, et al.
Veröffentlicht: (2024)
Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
von: Afonin, Nikita, et al.
Veröffentlicht: (2025)
von: Afonin, Nikita, et al.
Veröffentlicht: (2025)
MultiParaDetox: Extending Text Detoxification with Parallel Data to New Languages
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
von: Dementieva, Daryna, et al.
Veröffentlicht: (2024)
<think> So let's replace this phrase with insult... </think> Lessons learned from generation of toxic texts with LLMs
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
von: Pletenev, Sergey, et al.
Veröffentlicht: (2025)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
von: Murphy, Alexander, et al.
Veröffentlicht: (2025)
von: Murphy, Alexander, et al.
Veröffentlicht: (2025)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
von: Hase, Peter, et al.
Veröffentlicht: (2026)
von: Hase, Peter, et al.
Veröffentlicht: (2026)
Framework GNN-AID: Graph Neural Network Analysis Interpretation and Defense
von: Lukyanov, Kirill, et al.
Veröffentlicht: (2025)
von: Lukyanov, Kirill, et al.
Veröffentlicht: (2025)
Orientability of Causal Relations in Time Series using Summary Causal Graphs and Faithful Distributions
von: Loranchet, Timothée, et al.
Veröffentlicht: (2025)
von: Loranchet, Timothée, et al.
Veröffentlicht: (2025)
Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
von: Li, Jiachun, et al.
Veröffentlicht: (2024)
Breaking the Trade-Off Between Faithfulness and Expressiveness for Large Language Models
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
von: Yang, Chenxu, et al.
Veröffentlicht: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
von: Zaman, Kerem, et al.
Veröffentlicht: (2025)
Probabilistic Verification of Voice Anti-Spoofing Models
von: Kushnir, Evgeny, et al.
Veröffentlicht: (2026)
von: Kushnir, Evgeny, et al.
Veröffentlicht: (2026)
LLM-Guided Prompt Evolution for Password Guessing
von: Mazin, Vladimir A., et al.
Veröffentlicht: (2026)
von: Mazin, Vladimir A., et al.
Veröffentlicht: (2026)
Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
von: Young, Richard J.
Veröffentlicht: (2026)
von: Young, Richard J.
Veröffentlicht: (2026)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
von: Arcuschin, Iván, et al.
Veröffentlicht: (2025)
Equation discovery framework EPDE: Towards a better equation discovery
von: Maslyaev, Mikhail, et al.
Veröffentlicht: (2024)
von: Maslyaev, Mikhail, et al.
Veröffentlicht: (2024)
A Gap Between the Gaussian RKHS and Neural Networks: An Infinite-Center Asymptotic Analysis
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
Intermediate Languages Matter: Formal Choice Drives Neurosymbolic LLM Reasoning
von: Beiser, Alexander, et al.
Veröffentlicht: (2025)
von: Beiser, Alexander, et al.
Veröffentlicht: (2025)
GameUIAgent: An LLM-Powered Framework for Automated Game UI Design with Structured Intermediate Representation
von: Zeng, Wei, et al.
Veröffentlicht: (2026)
von: Zeng, Wei, et al.
Veröffentlicht: (2026)
When Bits Break Recourse: Counterfactual-Faithful Quantization
von: Yahyati, Chaymae, et al.
Veröffentlicht: (2026)
von: Yahyati, Chaymae, et al.
Veröffentlicht: (2026)
ReCo: Reliable Causal Chain Reasoning via Structural Causal Recurrent Neural Networks
von: Xiong, Kai, et al.
Veröffentlicht: (2022)
von: Xiong, Kai, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2025) -
Evolutionary Search for Automated Design of Uncertainty Quantification Methods
von: Seleznyov, Mikhail, et al.
Veröffentlicht: (2026) -
The benefits of query-based KGQA systems for complex and temporal questions in LLM era
von: Alekseev, Artem, et al.
Veröffentlicht: (2025) -
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026) -
Leveraging LLM Parametric Knowledge for Fact Checking without Retrieval
von: Vazhentsev, Artem, et al.
Veröffentlicht: (2026)