Counterfactuals As a Means for Evaluating Faithfulness of Attribution Methods in Autoregressive Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kamahi, Sepehr, Yaghoobzadeh, Yadollah |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GhazalBench: Usage-Grounded Evaluation of LLMs on Persian Ghazals
von: Kalhor, Ghazal, et al.
Veröffentlicht: (2026)
von: Kalhor, Ghazal, et al.
Veröffentlicht: (2026)
Layer-wise Positional Bias in Short-Context Language Modeling
von: Rahimi, Maryam, et al.
Veröffentlicht: (2026)
von: Rahimi, Maryam, et al.
Veröffentlicht: (2026)
Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
von: Rezaeimanesh, Sara, et al.
Veröffentlicht: (2024)
von: Rezaeimanesh, Sara, et al.
Veröffentlicht: (2024)
FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
von: Hosseini, Faezeh, et al.
Veröffentlicht: (2026)
von: Hosseini, Faezeh, et al.
Veröffentlicht: (2026)
Evaluating the Creativity of LLMs in Persian Literary Text Generation
von: Tourajmehr, Armin, et al.
Veröffentlicht: (2025)
von: Tourajmehr, Armin, et al.
Veröffentlicht: (2025)
Explanations of Large Language Models Explain Language Representations in the Brain
von: Rahimi, Maryam, et al.
Veröffentlicht: (2025)
von: Rahimi, Maryam, et al.
Veröffentlicht: (2025)
Benchmarking Large Language Models for Persian: A Preliminary Study Focusing on ChatGPT
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Abaskohi, Amirhossein, et al.
Veröffentlicht: (2024)
SOI Matters: Analyzing Multi-Setting Training Dynamics in Pretrained Language Models via Subsets of Interest
von: Vassef, Shayan, et al.
Veröffentlicht: (2025)
von: Vassef, Shayan, et al.
Veröffentlicht: (2025)
uTeBC-NLP at SemEval-2024 Task 9: Can LLMs be Lateral Thinkers?
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2024)
von: Sadeghi, Pouya, et al.
Veröffentlicht: (2024)
Comparative Study of Multilingual Idioms and Similes in Large Language Models
von: Khoshtab, Paria, et al.
Veröffentlicht: (2024)
von: Khoshtab, Paria, et al.
Veröffentlicht: (2024)
GenKnowSub: Improving Modularity and Reusability of LLMs through General Knowledge Subtraction
von: Bagherifard, Mohammadtaha, et al.
Veröffentlicht: (2025)
von: Bagherifard, Mohammadtaha, et al.
Veröffentlicht: (2025)
Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation
von: Sani, Samin Mahdizadeh, et al.
Veröffentlicht: (2024)
von: Sani, Samin Mahdizadeh, et al.
Veröffentlicht: (2024)
A Dual-Axis Taxonomy of Knowledge Editing for LLMs: From Mechanisms to Functions
von: Salehoof, Amir Mohammad, et al.
Veröffentlicht: (2025)
von: Salehoof, Amir Mohammad, et al.
Veröffentlicht: (2025)
PerCul: A Story-Driven Cultural Evaluation of LLMs in Persian
von: Monazzah, Erfan Moosavi, et al.
Veröffentlicht: (2025)
von: Monazzah, Erfan Moosavi, et al.
Veröffentlicht: (2025)
Synthia: Scalable Grounded Persona Generation from Social Media Data
von: Rahimzadeh, Vahid, et al.
Veröffentlicht: (2025)
von: Rahimzadeh, Vahid, et al.
Veröffentlicht: (2025)
Faithfulness vs. Safety: Evaluating LLM Behavior Under Counterfactual Medical Evidence
von: Mo, Kaijie, et al.
Veröffentlicht: (2026)
von: Mo, Kaijie, et al.
Veröffentlicht: (2026)
Enhancing Answer Attribution for Faithful Text Generation with Large Language Models
von: Vladika, Juraj, et al.
Veröffentlicht: (2024)
von: Vladika, Juraj, et al.
Veröffentlicht: (2024)
Faithfulness-QA: A Counterfactual Entity Substitution Dataset for Training Context-Faithful RAG Models
von: Ju, Li, et al.
Veröffentlicht: (2026)
von: Ju, Li, et al.
Veröffentlicht: (2026)
Counterfactual Simulation Training for Chain-of-Thought Faithfulness
von: Hase, Peter, et al.
Veröffentlicht: (2026)
von: Hase, Peter, et al.
Veröffentlicht: (2026)
Evaluating Counterfactual Strategic Reasoning in Large Language Models
von: Georgousis, Dimitrios, et al.
Veröffentlicht: (2026)
von: Georgousis, Dimitrios, et al.
Veröffentlicht: (2026)
Faithful Model Evaluation for Model-Based Metrics
von: Goyal, Palash, et al.
Veröffentlicht: (2023)
von: Goyal, Palash, et al.
Veröffentlicht: (2023)
Faithfulness Evaluation for Decoder-only LLM Attributions with Controlled Retained Information
von: Huang, Xin, et al.
Veröffentlicht: (2026)
von: Huang, Xin, et al.
Veröffentlicht: (2026)
RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
von: Han, Yunseok, et al.
Veröffentlicht: (2026)
von: Han, Yunseok, et al.
Veröffentlicht: (2026)
RSAT: Structured Attribution Makes Small Language Models Faithful Table Reasoners
von: Gajjar, Jugal, et al.
Veröffentlicht: (2026)
von: Gajjar, Jugal, et al.
Veröffentlicht: (2026)
Correctness is not Faithfulness in RAG Attributions
von: Wallat, Jonas, et al.
Veröffentlicht: (2024)
von: Wallat, Jonas, et al.
Veröffentlicht: (2024)
CounterBench: Evaluating and Improving Counterfactual Reasoning in Large Language Models
von: Chen, Yuefei, et al.
Veröffentlicht: (2025)
von: Chen, Yuefei, et al.
Veröffentlicht: (2025)
Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
von: Alon, Bar, et al.
Veröffentlicht: (2026)
von: Alon, Bar, et al.
Veröffentlicht: (2026)
Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
von: Jiang, Yifan, et al.
Veröffentlicht: (2026)
Teaching Language Models to Faithfully Express their Uncertainty
von: Eikema, Bryan, et al.
Veröffentlicht: (2025)
von: Eikema, Bryan, et al.
Veröffentlicht: (2025)
Incorporating Attribution Importance for Improving Faithfulness Metrics
von: Zhao, Zhixue, et al.
Veröffentlicht: (2023)
von: Zhao, Zhixue, et al.
Veröffentlicht: (2023)
Talent or Luck? Evaluating Attribution Bias in Large Language Models
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
von: Raj, Chahat, et al.
Veröffentlicht: (2025)
VALOR-EVAL: Holistic Coverage and Faithfulness Evaluation of Large Vision-Language Models
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
von: Qiu, Haoyi, et al.
Veröffentlicht: (2024)
Evaluating Reasoning Faithfulness in Medical Vision-Language Models using Multimodal Perturbations
von: Moll, Johannes, et al.
Veröffentlicht: (2025)
von: Moll, Johannes, et al.
Veröffentlicht: (2025)
Mapping Faithful Reasoning in Language Models
von: Li, Jiazheng, et al.
Veröffentlicht: (2025)
von: Li, Jiazheng, et al.
Veröffentlicht: (2025)
Faithfulness Measurable Masked Language Models
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
von: Madsen, Andreas, et al.
Veröffentlicht: (2023)
FaithUn: Toward Faithful Forgetting in Language Models by Investigating the Interconnectedness of Knowledge
von: Yang, Nakyeong, et al.
Veröffentlicht: (2025)
von: Yang, Nakyeong, et al.
Veröffentlicht: (2025)
Mitigating Large Language Model Hallucination with Faithful Finetuning
von: Hu, Minda, et al.
Veröffentlicht: (2024)
von: Hu, Minda, et al.
Veröffentlicht: (2024)
Faithfulness Metrics Don't Measure Faithfulness: A Meta-Evaluation with Ground Truth
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026)
von: Gur-Arieh, Yoav, et al.
Veröffentlicht: (2026)
Causal Autoregressive Diffusion Language Model
von: Ruan, Junhao, et al.
Veröffentlicht: (2026)
von: Ruan, Junhao, et al.
Veröffentlicht: (2026)
Reversal Invariance in Autoregressive Language Models
von: Sahasrabudhe, Mihir
Veröffentlicht: (2025)
von: Sahasrabudhe, Mihir
Veröffentlicht: (2025)
Ähnliche Einträge
-
GhazalBench: Usage-Grounded Evaluation of LLMs on Persian Ghazals
von: Kalhor, Ghazal, et al.
Veröffentlicht: (2026) -
Layer-wise Positional Bias in Short-Context Language Modeling
von: Rahimi, Maryam, et al.
Veröffentlicht: (2026) -
Large Language Models for Persian $ \leftrightarrow $ English Idiom Translation
von: Rezaeimanesh, Sara, et al.
Veröffentlicht: (2024) -
FFE-Hallu:Hallucinations in Fixed Figurative Expressions:Benchmark of Idioms and Proverbs in the Persian Language
von: Hosseini, Faezeh, et al.
Veröffentlicht: (2026) -
Evaluating the Creativity of LLMs in Persian Literary Text Generation
von: Tourajmehr, Armin, et al.
Veröffentlicht: (2025)