Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen, Van Bach, Seifert, Christin, Schlötterer, Jörg |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CEval: A Benchmark for Evaluating Counterfactual Text Generation
por: Nguyen, Van Bach, et al.
Publicado: (2024)
por: Nguyen, Van Bach, et al.
Publicado: (2024)
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
por: Nguyen, Van Bach, et al.
Publicado: (2024)
por: Nguyen, Van Bach, et al.
Publicado: (2024)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
por: Nguyen, Van Bach, et al.
Publicado: (2022)
por: Nguyen, Van Bach, et al.
Publicado: (2022)
Persuasion Tokens for Editing Factual Knowledge in LLMs
por: Youssef, Paul, et al.
Publicado: (2026)
por: Youssef, Paul, et al.
Publicado: (2026)
Tracing and Reversing Edits in LLMs
por: Youssef, Paul, et al.
Publicado: (2025)
por: Youssef, Paul, et al.
Publicado: (2025)
How to Make LLMs Forget: On Reversing In-Context Knowledge Edits
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
Enhancing Fact Retrieval in PLMs through Truthfulness
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
A Second Look on BASS -- Boosting Abstractive Summarization with Unified Semantic Graphs -- A Replication Study
por: Koraş, Osman Alperen, et al.
Publicado: (2024)
por: Koraş, Osman Alperen, et al.
Publicado: (2024)
Has this Fact been Edited? Detecting Knowledge Edits in Language Models
por: Youssef, Paul, et al.
Publicado: (2024)
por: Youssef, Paul, et al.
Publicado: (2024)
Behavioral Analysis of Information Salience in Large Language Models
por: Trienes, Jan, et al.
Publicado: (2025)
por: Trienes, Jan, et al.
Publicado: (2025)
Position: Editing Large Language Models Poses Serious Safety Risks
por: Youssef, Paul, et al.
Publicado: (2025)
por: Youssef, Paul, et al.
Publicado: (2025)
Can Fine-Tuning Erase Your Edits? On the Fragile Coexistence of Knowledge Editing and Adaptation
por: Cheng, Yinjie, et al.
Publicado: (2025)
por: Cheng, Yinjie, et al.
Publicado: (2025)
Marcel: A Lightweight and Open-Source Conversational Agent for University Student Support
por: Trienes, Jan, et al.
Publicado: (2025)
por: Trienes, Jan, et al.
Publicado: (2025)
Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
por: Wang, Qianli, et al.
Publicado: (2025)
por: Wang, Qianli, et al.
Publicado: (2025)
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
por: Wang, Qianli, et al.
Publicado: (2026)
por: Wang, Qianli, et al.
Publicado: (2026)
InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
por: Trienes, Jan, et al.
Publicado: (2024)
por: Trienes, Jan, et al.
Publicado: (2024)
An XAI-based Analysis of Shortcut Learning in Neural Networks
por: Le, Phuong Quynh, et al.
Publicado: (2025)
por: Le, Phuong Quynh, et al.
Publicado: (2025)
Is Last Layer Re-Training Truly Sufficient for Robustness to Spurious Correlations?
por: Le, Phuong Quynh, et al.
Publicado: (2023)
por: Le, Phuong Quynh, et al.
Publicado: (2023)
Explanation format does not matter; but explanations do -- An Eggsbert study on explaining Bayesian Optimisation tasks
por: Chakraborty, Tanmay, et al.
Publicado: (2025)
por: Chakraborty, Tanmay, et al.
Publicado: (2025)
DRIV-EX: Counterfactual Explanations for Driving LLMs
por: Cardiel, Amaia, et al.
Publicado: (2026)
por: Cardiel, Amaia, et al.
Publicado: (2026)
Towards Faithful Explanations for Text Classification with Robustness Improvement and Explanation Guided Training
por: Li, Dongfang, et al.
Publicado: (2023)
por: Li, Dongfang, et al.
Publicado: (2023)
Towards Interpretable Deep Neural Networks for Tabular Data
por: Elhadri, Khawla, et al.
Publicado: (2025)
por: Elhadri, Khawla, et al.
Publicado: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
por: Elhadri, Khawla, et al.
Publicado: (2025)
por: Elhadri, Khawla, et al.
Publicado: (2025)
Investigating the Impact of Randomness on Reproducibility in Computer Vision: A Study on Applications in Civil Engineering and Medicine
por: Eryılmaz, Bahadır, et al.
Publicado: (2024)
por: Eryılmaz, Bahadır, et al.
Publicado: (2024)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
por: Hamman, Faisal, et al.
Publicado: (2025)
por: Hamman, Faisal, et al.
Publicado: (2025)
Invariant Learning with Annotation-free Environments
por: Le, Phuong Quynh, et al.
Publicado: (2025)
por: Le, Phuong Quynh, et al.
Publicado: (2025)
Out of Spuriousity: Improving Robustness to Spurious Correlations without Group Annotations
por: Le, Phuong Quynh, et al.
Publicado: (2024)
por: Le, Phuong Quynh, et al.
Publicado: (2024)
A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers
por: McAleese, Stephen, et al.
Publicado: (2024)
por: McAleese, Stephen, et al.
Publicado: (2024)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
por: Limpijankit, Marvin, et al.
Publicado: (2025)
por: Limpijankit, Marvin, et al.
Publicado: (2025)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
por: Mayne, Harry, et al.
Publicado: (2025)
por: Mayne, Harry, et al.
Publicado: (2025)
Mitigating Text Toxicity with Counterfactual Generation
por: Bhan, Milan, et al.
Publicado: (2024)
por: Bhan, Milan, et al.
Publicado: (2024)
WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
por: Jiang, Yiwen, et al.
Publicado: (2025)
por: Jiang, Yiwen, et al.
Publicado: (2025)
Prompt-Counterfactual Explanations for Generative AI System Behavior
por: Goethals, Sofie, et al.
Publicado: (2026)
por: Goethals, Sofie, et al.
Publicado: (2026)
Funzac at CoMeDi Shared Task: Modeling Annotator Disagreement from Word-In-Context Perspectives
por: Sarumi, Olufunke O., et al.
Publicado: (2025)
por: Sarumi, Olufunke O., et al.
Publicado: (2025)
The Impact of Annotator Personas on LLM Behavior Across the Perspectivism Spectrum
por: Sarumi, Olufunke O., et al.
Publicado: (2025)
por: Sarumi, Olufunke O., et al.
Publicado: (2025)
Patch-based Intuitive Multimodal Prototypes Network (PIMPNet) for Alzheimer's Disease classification
por: De Santi, Lisa Anita, et al.
Publicado: (2024)
por: De Santi, Lisa Anita, et al.
Publicado: (2024)
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
por: Toker, Gilat, et al.
Publicado: (2026)
por: Toker, Gilat, et al.
Publicado: (2026)
A Combined Encoder and Transformer Approach for Coherent and High-Quality Text Generation
por: Chen, Jiajing, et al.
Publicado: (2024)
por: Chen, Jiajing, et al.
Publicado: (2024)
Can LLMs Generate High-Quality Task-Specific Conversations?
por: Li, Shengqi, et al.
Publicado: (2025)
por: Li, Shengqi, et al.
Publicado: (2025)
Ejemplares similares
-
CEval: A Benchmark for Evaluating Counterfactual Text Generation
por: Nguyen, Van Bach, et al.
Publicado: (2024) -
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
por: Nguyen, Van Bach, et al.
Publicado: (2024) -
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
por: Nguyen, Van Bach, et al.
Publicado: (2022) -
Persuasion Tokens for Editing Factual Knowledge in LLMs
por: Youssef, Paul, et al.
Publicado: (2026) -
Tracing and Reversing Edits in LLMs
por: Youssef, Paul, et al.
Publicado: (2025)