LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
Fuente:
arXiv
Salvato in:
| Autori principali: | Nguyen, Van Bach, Youssef, Paul, Seifert, Christin, Schlötterer, Jörg |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
CEval: A Benchmark for Evaluating Counterfactual Text Generation
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024)
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
di: Nguyen, Van Bach, et al.
Pubblicazione: (2022)
di: Nguyen, Van Bach, et al.
Pubblicazione: (2022)
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
di: Nguyen, Van Bach, et al.
Pubblicazione: (2025)
di: Nguyen, Van Bach, et al.
Pubblicazione: (2025)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
di: Cheng, Yinjie, et al.
Pubblicazione: (2025)
di: Cheng, Yinjie, et al.
Pubblicazione: (2025)
Persuasion Tokens for Editing Factual Knowledge in LLMs
di: Youssef, Paul, et al.
Pubblicazione: (2026)
di: Youssef, Paul, et al.
Pubblicazione: (2026)
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
di: Wang, Qianli, et al.
Pubblicazione: (2026)
di: Wang, Qianli, et al.
Pubblicazione: (2026)
Tracing and Reversing Edits in LLMs
di: Youssef, Paul, et al.
Pubblicazione: (2025)
di: Youssef, Paul, et al.
Pubblicazione: (2025)
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
Enhancing Fact Retrieval in PLMs through Truthfulness
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
di: Youssef, Paul, et al.
Pubblicazione: (2024)
di: Youssef, Paul, et al.
Pubblicazione: (2024)
Position: Editing Large Language Models Poses Serious Safety Risks
di: Youssef, Paul, et al.
Pubblicazione: (2025)
di: Youssef, Paul, et al.
Pubblicazione: (2025)
Investigating the Impact of Randomness on Reproducibility in Computer Vision: A Study on Applications in Civil Engineering and Medicine
di: Eryılmaz, Bahadır, et al.
Pubblicazione: (2024)
di: Eryılmaz, Bahadır, et al.
Pubblicazione: (2024)
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
di: Koraş, Osman Alperen, et al.
Pubblicazione: (2024)
di: Koraş, Osman Alperen, et al.
Pubblicazione: (2024)
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
di: Le, Phuong Quynh, et al.
Pubblicazione: (2024)
di: Le, Phuong Quynh, et al.
Pubblicazione: (2024)
Behavioral Analysis of Information Salience in Large Language Models
di: Trienes, Jan, et al.
Pubblicazione: (2025)
di: Trienes, Jan, et al.
Pubblicazione: (2025)
Comprehensive Study on German Language Models for Clinical and Biomedical Text Understanding
di: Idrissi-Yaghir, Ahmad, et al.
Pubblicazione: (2024)
di: Idrissi-Yaghir, Ahmad, et al.
Pubblicazione: (2024)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
di: Ghosh, Rajarshi, et al.
Pubblicazione: (2025)
di: Ghosh, Rajarshi, et al.
Pubblicazione: (2025)
Comprehensive and Practical Evaluation of Retrieval-Augmented Generation Systems for Medical Question Answering
di: Ngo, Nghia Trung, et al.
Pubblicazione: (2024)
di: Ngo, Nghia Trung, et al.
Pubblicazione: (2024)
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
di: Wang, Qianli, et al.
Pubblicazione: (2025)
di: Wang, Qianli, et al.
Pubblicazione: (2025)
A Comprehensive Evaluation of Cognitive Biases in LLMs
di: Malberg, Simon, et al.
Pubblicazione: (2024)
di: Malberg, Simon, et al.
Pubblicazione: (2024)
Can LLMs Explain Themselves Counterfactually?
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2025)
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2025)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
di: Jin, Keyan, et al.
Pubblicazione: (2025)
di: Jin, Keyan, et al.
Pubblicazione: (2025)
Marcel: A Lightweight and Open-Source Conversational Agent for University Student Support
di: Trienes, Jan, et al.
Pubblicazione: (2025)
di: Trienes, Jan, et al.
Pubblicazione: (2025)
ViLLM-Eval: A Comprehensive Evaluation Suite for Vietnamese Large Language Models
di: Nguyen, Trong-Hieu, et al.
Pubblicazione: (2024)
di: Nguyen, Trong-Hieu, et al.
Pubblicazione: (2024)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
di: Turk, Matt
Pubblicazione: (2026)
di: Turk, Matt
Pubblicazione: (2026)
DSC2025 -- ViHallu Challenge: Detecting Hallucination in Vietnamese LLMs
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2026)
di: Nguyen, Anh Thi-Hoang, et al.
Pubblicazione: (2026)
Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in Large Language Models
di: Lewis, Martha, et al.
Pubblicazione: (2024)
di: Lewis, Martha, et al.
Pubblicazione: (2024)
Zero-shot LLM-guided Counterfactual Generation: A Case Study on NLP Model Evaluation
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
di: Bhattacharjee, Amrita, et al.
Pubblicazione: (2024)
Facilitating Cognitive Accessibility with LLMs: A Multi-Task Approach to Easy-to-Read Text Generation
di: Ledoyen, François, et al.
Pubblicazione: (2025)
di: Ledoyen, François, et al.
Pubblicazione: (2025)
EasyJudge: an Easy-to-use Tool for Comprehensive Response Evaluation of LLMs
di: Li, Yijie, et al.
Pubblicazione: (2024)
di: Li, Yijie, et al.
Pubblicazione: (2024)
Does Using Counterfactual Help LLMs Explain Textual Importance in Classification?
di: Tan, Nelvin, et al.
Pubblicazione: (2025)
di: Tan, Nelvin, et al.
Pubblicazione: (2025)
Can LLMs Reason About Program Semantics? A Comprehensive Evaluation of LLMs on Formal Specification Inference
di: Le-Cong, Thanh, et al.
Pubblicazione: (2025)
di: Le-Cong, Thanh, et al.
Pubblicazione: (2025)
Deriving Strategic Market Insights with Large Language Models: A Benchmark for Forward Counterfactual Generation
di: Ong, Keane, et al.
Pubblicazione: (2025)
di: Ong, Keane, et al.
Pubblicazione: (2025)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
di: Toker, Gilat, et al.
Pubblicazione: (2026)
di: Toker, Gilat, et al.
Pubblicazione: (2026)
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
di: Azmat, Muneeza, et al.
Pubblicazione: (2025)
MedEthicsQA: A Comprehensive Question Answering Benchmark for Medical Ethics Evaluation of LLMs
di: Wei, Jianhui, et al.
Pubblicazione: (2025)
di: Wei, Jianhui, et al.
Pubblicazione: (2025)
Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach
di: Li, Zhuowan, et al.
Pubblicazione: (2024)
di: Li, Zhuowan, et al.
Pubblicazione: (2024)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
SCENE: Evaluating Explainable AI Techniques Using Soft Counterfactuals
di: Zheng, Haoran, et al.
Pubblicazione: (2024)
di: Zheng, Haoran, et al.
Pubblicazione: (2024)
Documenti analoghi
-
CEval: A Benchmark for Evaluating Counterfactual Text Generation
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024) -
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
di: Nguyen, Van Bach, et al.
Pubblicazione: (2022) -
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
di: Nguyen, Van Bach, et al.
Pubblicazione: (2025) -
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
di: Youssef, Paul, et al.
Pubblicazione: (2024) -
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
di: Cheng, Yinjie, et al.
Pubblicazione: (2025)