LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Van Bach, Youssef, Paul, Seifert, Christin, Schlötterer, Jörg |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CEval: A Benchmark for Evaluating Counterfactual Text Generation
by: Nguyen, Van Bach, et al.
Published: (2024)
by: Nguyen, Van Bach, et al.
Published: (2024)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
by: Nguyen, Van Bach, et al.
Published: (2022)
by: Nguyen, Van Bach, et al.
Published: (2022)
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
by: Nguyen, Van Bach, et al.
Published: (2025)
by: Nguyen, Van Bach, et al.
Published: (2025)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
by: Cheng, Yinjie, et al.
Published: (2025)
by: Cheng, Yinjie, et al.
Published: (2025)
Persuasion Tokens for Editing Factual Knowledge in LLMs
by: Youssef, Paul, et al.
Published: (2026)
by: Youssef, Paul, et al.
Published: (2026)
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
by: Wang, Qianli, et al.
Published: (2026)
by: Wang, Qianli, et al.
Published: (2026)
Tracing and Reversing Edits in LLMs
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Enhancing Fact Retrieval in PLMs through Truthfulness
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
by: Youssef, Paul, et al.
Published: (2024)
by: Youssef, Paul, et al.
Published: (2024)
Position: Editing Large Language Models Poses Serious Safety Risks
by: Youssef, Paul, et al.
Published: (2025)
by: Youssef, Paul, et al.
Published: (2025)
Investigating the Impact of Randomness on Reproducibility in Computer Vision: A Study on Applications in Civil Engineering and Medicine
by: Eryılmaz, Bahadır, et al.
Published: (2024)
by: Eryılmaz, Bahadır, et al.
Published: (2024)
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
by: Koraş, Osman Alperen, et al.
Published: (2024)
by: Koraş, Osman Alperen, et al.
Published: (2024)
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
by: Le, Phuong Quynh, et al.
Published: (2024)
by: Le, Phuong Quynh, et al.
Published: (2024)
Behavioral Analysis of Information Salience in Large Language Models
by: Trienes, Jan, et al.
Published: (2025)
by: Trienes, Jan, et al.
Published: (2025)
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
by: Idrissi-Yaghir, Ahmad, et al.
Published: (2024)
by: Idrissi-Yaghir, Ahmad, et al.
Published: (2024)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
by: Ghosh, Rajarshi, et al.
Published: (2025)
by: Ghosh, Rajarshi, et al.
Published: (2025)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
by: Ngo, Nghia Trung, et al.
Published: (2024)
by: Ngo, Nghia Trung, et al.
Published: (2024)
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
by: Wang, Qianli, et al.
Published: (2025)
by: Wang, Qianli, et al.
Published: (2025)
A Comprehensive Evaluation of Cognitive Biases in LLMs
by: Malberg, Simon, et al.
Published: (2024)
by: Malberg, Simon, et al.
Published: (2024)
Can LLMs Explain Themselves Counterfactually?
by: Dehghanighobadi, Zahra, et al.
Published: (2025)
by: Dehghanighobadi, Zahra, et al.
Published: (2025)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
by: Jin, Keyan, et al.
Published: (2025)
by: Jin, Keyan, et al.
Published: (2025)
Marcel: A Lightweight and Open-Source Conversational Agent for University Student Support
by: Trienes, Jan, et al.
Published: (2025)
by: Trienes, Jan, et al.
Published: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
by: Nguyen, Trong-Hieu, et al.
Published: (2024)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
by: Nguyen, Anh Thi-Hoang, et al.
Published: (2026)
Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in Large Language Models
by: Lewis, Martha, et al.
Published: (2024)
by: Lewis, Martha, et al.
Published: (2024)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
by: Bhattacharjee, Amrita, et al.
Published: (2024)
by: Bhattacharjee, Amrita, et al.
Published: (2024)
Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
by: Ledoyen, François, et al.
Published: (2025)
by: Ledoyen, François, et al.
Published: (2025)
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs
by: Li, Yijie, et al.
Published: (2024)
by: Li, Yijie, et al.
Published: (2024)
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
by: Tan, Nelvin, et al.
Published: (2025)
by: Tan, Nelvin, et al.
Published: (2025)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
by: Le-Cong, Thanh, et al.
Published: (2025)
by: Le-Cong, Thanh, et al.
Published: (2025)
Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation
by: Ong, Keane, et al.
Published: (2025)
by: Ong, Keane, et al.
Published: (2025)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
by: Toker, Gilat, et al.
Published: (2026)
by: Toker, Gilat, et al.
Published: (2026)
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
by: Azmat, Muneeza, et al.
Published: (2025)
by: Azmat, Muneeza, et al.
Published: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
by: Wei, Jianhui, et al.
Published: (2025)
by: Wei, Jianhui, et al.
Published: (2025)
Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
by: Li, Zhuowan, et al.
Published: (2024)
by: Li, Zhuowan, et al.
Published: (2024)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
by: Limpijankit, Marvin, et al.
Published: (2025)
by: Limpijankit, Marvin, et al.
Published: (2025)
SCENE: Evaluating Explainable AI Techniques Using Soft Counterfactuals
by: Zheng, Haoran, et al.
Published: (2024)
by: Zheng, Haoran, et al.
Published: (2024)
Similar Items
-
CEval: A Benchmark for Evaluating Counterfactual Text Generation
by: Nguyen, Van Bach, et al.
Published: (2024) -
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
by: Nguyen, Van Bach, et al.
Published: (2022) -
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
by: Nguyen, Van Bach, et al.
Published: (2025) -
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
by: Youssef, Paul, et al.
Published: (2024) -
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
by: Cheng, Yinjie, et al.
Published: (2025)