A Causal Lens for Evaluating Faithfulness Metrics
Fuente:
arXiv
Guardado en:
| Autores principales: | Zaman, Kerem, Srivastava, Shashank |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
por: Zaman, Kerem, et al.
Publicado: (2025)
por: Zaman, Kerem, et al.
Publicado: (2025)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
por: Zaman, Kerem, et al.
Publicado: (2023)
por: Zaman, Kerem, et al.
Publicado: (2023)
RCT Rejection Sampling for Causal Estimation Evaluation
por: Keith, Katherine A., et al.
Publicado: (2023)
por: Keith, Katherine A., et al.
Publicado: (2023)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
por: Cai, Hengrui, et al.
Publicado: (2023)
por: Cai, Hengrui, et al.
Publicado: (2023)
Text Rationalization for Robust Causal Effect Estimation
por: Zhang, Lijinghua, et al.
Publicado: (2025)
por: Zhang, Lijinghua, et al.
Publicado: (2025)
ALCM: Autonomous LLM-Augmented Causal Discovery Framework
por: Khatibi, Elahe, et al.
Publicado: (2024)
por: Khatibi, Elahe, et al.
Publicado: (2024)
CLEAR: Can Language Models Really Understand Causal Graphs?
por: Chen, Sirui, et al.
Publicado: (2024)
por: Chen, Sirui, et al.
Publicado: (2024)
A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
por: Sarmah, Bhaskarjit, et al.
Publicado: (2024)
por: Sarmah, Bhaskarjit, et al.
Publicado: (2024)
Industrial-Grade Smart Troubleshooting through Causal Technical Language Processing: a Proof of Concept
por: Trilla, Alexandre, et al.
Publicado: (2024)
por: Trilla, Alexandre, et al.
Publicado: (2024)
Language Models as Causal Effect Generators
por: Bynum, Lucius E. J., et al.
Publicado: (2024)
por: Bynum, Lucius E. J., et al.
Publicado: (2024)
Rating Multi-Modal Time-Series Forecasting Models (MM-TSFM) for Robustness Through a Causal Lens
por: Lakkaraju, Kausik, et al.
Publicado: (2024)
por: Lakkaraju, Kausik, et al.
Publicado: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
por: Kasetty, Tejas, et al.
Publicado: (2024)
por: Kasetty, Tejas, et al.
Publicado: (2024)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
por: Boyeau, Pierre, et al.
Publicado: (2024)
por: Boyeau, Pierre, et al.
Publicado: (2024)
Transformer Circuit Faithfulness Metrics are not Robust
por: Miller, Joseph, et al.
Publicado: (2024)
por: Miller, Joseph, et al.
Publicado: (2024)
(Mis)Fitting: A Survey of Scaling Laws
por: Li, Margaret, et al.
Publicado: (2025)
por: Li, Margaret, et al.
Publicado: (2025)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
por: Chew, Robert, et al.
Publicado: (2026)
por: Chew, Robert, et al.
Publicado: (2026)
Incorporating Attribution Importance for Improving Faithfulness Metrics
por: Zhao, Zhixue, et al.
Publicado: (2023)
por: Zhao, Zhixue, et al.
Publicado: (2023)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
por: Kıcıman, Emre, et al.
Publicado: (2023)
por: Kıcıman, Emre, et al.
Publicado: (2023)
The Leaderboard Illusion
por: Singh, Shivalika, et al.
Publicado: (2025)
por: Singh, Shivalika, et al.
Publicado: (2025)
Majority of the Bests: Improving Best-of-N via Bootstrapping
por: Rakhsha, Amin, et al.
Publicado: (2025)
por: Rakhsha, Amin, et al.
Publicado: (2025)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
por: Bhardwaj, Dhrupad, et al.
Publicado: (2025)
por: Bhardwaj, Dhrupad, et al.
Publicado: (2025)
Efficient Exploration for LLMs
por: Dwaracherla, Vikranth, et al.
Publicado: (2024)
por: Dwaracherla, Vikranth, et al.
Publicado: (2024)
Adaptive Uncertainty Quantification for Generative AI
por: Kim, Jungeum, et al.
Publicado: (2024)
por: Kim, Jungeum, et al.
Publicado: (2024)
Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks
por: Hua, Wenyue, et al.
Publicado: (2024)
por: Hua, Wenyue, et al.
Publicado: (2024)
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
por: Do, Heejin, et al.
Publicado: (2026)
por: Do, Heejin, et al.
Publicado: (2026)
A Causal Framework for Evaluating ICU Discharge Strategies
por: Simha, Sagar Nagaraj, et al.
Publicado: (2026)
por: Simha, Sagar Nagaraj, et al.
Publicado: (2026)
CausalCompass: Evaluating the Robustness of Time-Series Causal Discovery in Misspecified Scenarios
por: Yi, Huiyang, et al.
Publicado: (2026)
por: Yi, Huiyang, et al.
Publicado: (2026)
Removing Spurious Correlation from Neural Network Interpretations
por: Fotouhi, Milad, et al.
Publicado: (2024)
por: Fotouhi, Milad, et al.
Publicado: (2024)
Black Box Causal Inference: Effect Estimation via Meta Prediction
por: Bynum, Lucius E. J., et al.
Publicado: (2025)
por: Bynum, Lucius E. J., et al.
Publicado: (2025)
MarginSel : Max-Margin Demonstration Selection for LLMs
por: Ambati, Rajeev Bhatt, et al.
Publicado: (2025)
por: Ambati, Rajeev Bhatt, et al.
Publicado: (2025)
Evaluation of Stress Detection as Time Series Events -- A Novel Window-Based F1-Metric
por: Skat-Rørdam, Harald Vilhelm, et al.
Publicado: (2025)
por: Skat-Rørdam, Harald Vilhelm, et al.
Publicado: (2025)
Bridging the Unavoidable A Priori: A Framework for Comparative Causal Modeling
por: Hovmand, Peter S., et al.
Publicado: (2025)
por: Hovmand, Peter S., et al.
Publicado: (2025)
Causal Fairness under Unobserved Confounding: A Neural Sensitivity Framework
por: Schröder, Maresa, et al.
Publicado: (2023)
por: Schröder, Maresa, et al.
Publicado: (2023)
Identifying Causal Effects Under Functional Dependencies
por: Chen, Yizuo, et al.
Publicado: (2024)
por: Chen, Yizuo, et al.
Publicado: (2024)
CARL-GT: Evaluating Causal Reasoning Capabilities of Large Language Models
por: Tu, Ruibo, et al.
Publicado: (2024)
por: Tu, Ruibo, et al.
Publicado: (2024)
Towards Robust Evaluation: A Comprehensive Taxonomy of Datasets and Metrics for Open Domain Question Answering in the Era of Large Language Models
por: Srivastava, Akchay, et al.
Publicado: (2024)
por: Srivastava, Akchay, et al.
Publicado: (2024)
Integrating Large Language Models in Causal Discovery: A Statistical Causal Approach
por: Takayama, Masayuki, et al.
Publicado: (2024)
por: Takayama, Masayuki, et al.
Publicado: (2024)
Multi-Domain Causal Discovery in Bijective Causal Models
por: Jalaldoust, Kasra, et al.
Publicado: (2025)
por: Jalaldoust, Kasra, et al.
Publicado: (2025)
Learning Causal Abstractions of Linear Structural Causal Models
por: Massidda, Riccardo, et al.
Publicado: (2024)
por: Massidda, Riccardo, et al.
Publicado: (2024)
LIDS: LLM Summary Inference Under the Layered Lens
por: Park, Dylan, et al.
Publicado: (2026)
por: Park, Dylan, et al.
Publicado: (2026)
Ejemplares similares
-
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
por: Zaman, Kerem, et al.
Publicado: (2025) -
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
por: Zaman, Kerem, et al.
Publicado: (2023) -
RCT Rejection Sampling for Causal Estimation Evaluation
por: Keith, Katherine A., et al.
Publicado: (2023) -
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
por: Cai, Hengrui, et al.
Publicado: (2023) -
Text Rationalization for Robust Causal Effect Estimation
por: Zhang, Lijinghua, et al.
Publicado: (2025)