FaithCoT-Bench: Benchmarking Instance-Level Faithfulness of Chain-of-Thought Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Xu, Wang, Song, Tan, Zhen, Yao, Laura, Zhao, Xinyu, Xu, Kaidi, Wang, Xin, Chen, Tianlong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
di: Lv, Weijiang, et al.
Pubblicazione: (2026)
di: Lv, Weijiang, et al.
Pubblicazione: (2026)
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
di: Mittal, Avni, et al.
Pubblicazione: (2026)
di: Mittal, Avni, et al.
Pubblicazione: (2026)
Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy
di: Shen, Xu, et al.
Pubblicazione: (2026)
di: Shen, Xu, et al.
Pubblicazione: (2026)
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
di: Lv, Weijiang, et al.
Pubblicazione: (2026)
di: Lv, Weijiang, et al.
Pubblicazione: (2026)
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)
Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder
di: Xu, Zhen, et al.
Pubblicazione: (2025)
di: Xu, Zhen, et al.
Pubblicazione: (2025)
Towards Better Chain-of-Thought: A Reflection on Effectiveness and Faithfulness
di: Li, Jiachun, et al.
Pubblicazione: (2024)
di: Li, Jiachun, et al.
Pubblicazione: (2024)
Lie to Me: How Faithful Is Chain-of-Thought Reasoning in Reasoning Models?
di: Young, Richard J.
Pubblicazione: (2026)
di: Young, Richard J.
Pubblicazione: (2026)
Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning
di: Parekh, Swapnil
Pubblicazione: (2026)
di: Parekh, Swapnil
Pubblicazione: (2026)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
di: Hase, Peter, et al.
Pubblicazione: (2026)
di: Hase, Peter, et al.
Pubblicazione: (2026)
Measuring Chain-of-Thought Monitorability Through Faithfulness and Verbosity
di: Meek, Austin, et al.
Pubblicazione: (2025)
di: Meek, Austin, et al.
Pubblicazione: (2025)
Faithful Logical Reasoning via Symbolic Chain-of-Thought
di: Xu, Jundong, et al.
Pubblicazione: (2024)
di: Xu, Jundong, et al.
Pubblicazione: (2024)
Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning
di: Jia, Jinghan, et al.
Pubblicazione: (2026)
di: Jia, Jinghan, et al.
Pubblicazione: (2026)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
di: Zaman, Kerem, et al.
Pubblicazione: (2025)
di: Zaman, Kerem, et al.
Pubblicazione: (2025)
VinaBench: Benchmark for Faithful and Consistent Visual Narratives
di: Gao, Silin, et al.
Pubblicazione: (2025)
di: Gao, Silin, et al.
Pubblicazione: (2025)
FRIT: Using Causal Importance to Improve Chain-of-Thought Faithfulness
di: Swaroop, Anand, et al.
Pubblicazione: (2025)
di: Swaroop, Anand, et al.
Pubblicazione: (2025)
FloCA: Towards Faithful and Logically Consistent Flowchart Reasoning
di: Zou, Jinzi, et al.
Pubblicazione: (2026)
di: Zou, Jinzi, et al.
Pubblicazione: (2026)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
di: Sun, Jingyi, et al.
Pubblicazione: (2026)
di: Sun, Jingyi, et al.
Pubblicazione: (2026)
Faithful or Just Plausible? Evaluating the Faithfulness of Closed-Source LLMs in Medical Reasoning
di: Afolabi, Halimat, et al.
Pubblicazione: (2026)
di: Afolabi, Halimat, et al.
Pubblicazione: (2026)
The Quest for Efficient Reasoning: A Data-Centric Benchmark to CoT Distillation
di: Zhang, Ruichen, et al.
Pubblicazione: (2025)
di: Zhang, Ruichen, et al.
Pubblicazione: (2025)
Exploring Knowledge Conflicts for Faithful LLM Reasoning: Benchmark and Method
di: Zhao, Tianzhe, et al.
Pubblicazione: (2026)
di: Zhao, Tianzhe, et al.
Pubblicazione: (2026)
FaithBench: A Diverse Hallucination Benchmark for Summarization by Modern LLMs
di: Bao, Forrest Sheng, et al.
Pubblicazione: (2024)
di: Bao, Forrest Sheng, et al.
Pubblicazione: (2024)
Faithful Density-Peaks Clustering via Matrix Computations on MPI Parallelization System
di: Xu, Ji, et al.
Pubblicazione: (2024)
di: Xu, Ji, et al.
Pubblicazione: (2024)
FaithLens: Detecting and Explaining Faithfulness Hallucination
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
di: Si, Shuzheng, et al.
Pubblicazione: (2025)
Mechanistic Evidence for Faithfulness Decay in Chain-of-Thought Reasoning
di: Ye, Donald, et al.
Pubblicazione: (2026)
di: Ye, Donald, et al.
Pubblicazione: (2026)
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
di: Tutek, Martin, et al.
Pubblicazione: (2025)
di: Tutek, Martin, et al.
Pubblicazione: (2025)
Stake the Points: Structure-Faithful Instance Unlearning
di: Hong, Kiseong, et al.
Pubblicazione: (2026)
di: Hong, Kiseong, et al.
Pubblicazione: (2026)
PaLMR: Towards Faithful Visual Reasoning via Multimodal Process Alignment
di: Li, Yantao, et al.
Pubblicazione: (2026)
di: Li, Yantao, et al.
Pubblicazione: (2026)
Faithful-First Reasoning, Planning, and Acting for Multimodal LLMs
di: Li, Junxian, et al.
Pubblicazione: (2025)
di: Li, Junxian, et al.
Pubblicazione: (2025)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
di: Dai, Yang, et al.
Pubblicazione: (2026)
di: Dai, Yang, et al.
Pubblicazione: (2026)
Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
di: Tamber, Manveer Singh, et al.
Pubblicazione: (2025)
Understanding Reasoning in Chain-of-Thought from the Hopfieldian View
di: Hu, Lijie, et al.
Pubblicazione: (2024)
di: Hu, Lijie, et al.
Pubblicazione: (2024)
Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
di: Paul, Debjit, et al.
Pubblicazione: (2024)
di: Paul, Debjit, et al.
Pubblicazione: (2024)
Evaluating GRPO and DPO for Faithful Chain-of-Thought Reasoning in LLMs
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
di: Mohammadi, Hadi, et al.
Pubblicazione: (2025)
On the Hardness of Faithful Chain-of-Thought Reasoning in Large Language Models
di: Tanneru, Sree Harsha, et al.
Pubblicazione: (2024)
di: Tanneru, Sree Harsha, et al.
Pubblicazione: (2024)
DatBench: Discriminative, Faithful, and Efficient VLM Evaluations
di: DatologyAI, et al.
Pubblicazione: (2026)
di: DatologyAI, et al.
Pubblicazione: (2026)
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
di: Han, Yunseok, et al.
Pubblicazione: (2026)
di: Han, Yunseok, et al.
Pubblicazione: (2026)
FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
di: Wang, Yuanzhi, et al.
Pubblicazione: (2026)
di: Wang, Yuanzhi, et al.
Pubblicazione: (2026)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
di: Balasubramanian, Sriram, et al.
Pubblicazione: (2025)
Measuring Faithfulness Depends on How You Measure: Classifier Sensitivity in LLM Chain-of-Thought Evaluation
di: Young, Richard J.
Pubblicazione: (2026)
di: Young, Richard J.
Pubblicazione: (2026)
Documenti analoghi
-
SPD-Faith Bench: Diagnosing and Improving Faithfulness in Chain-of-Thought for Multimodal Large Language Models
di: Lv, Weijiang, et al.
Pubblicazione: (2026) -
C2-Faith: Benchmarking LLM Judges for Causal and Coverage Faithfulness in Chain-of-Thought Reasoning
di: Mittal, Avni, et al.
Pubblicazione: (2026) -
Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy
di: Shen, Xu, et al.
Pubblicazione: (2026) -
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought
di: Lv, Weijiang, et al.
Pubblicazione: (2026) -
Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
di: Arcuschin, Iván, et al.
Pubblicazione: (2025)