Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Qianli, Nguyen, Van Bach, Feldhus, Nils, Villa-Arenas, Luis Felipe, Seifert, Christin, Möller, Sebastian, Schmitt, Vera |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation
von: Wang, Qianli, et al.
Veröffentlicht: (2025)
von: Wang, Qianli, et al.
Veröffentlicht: (2025)
iFlip: Iterative Feedback-driven Counterfactual Example Refinement
von: Wang, Yilong, et al.
Veröffentlicht: (2026)
von: Wang, Yilong, et al.
Veröffentlicht: (2026)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2024)
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2024)
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2024)
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2024)
Cross-Refine: Improving Natural Language Explanation Generation by Learning in Tandem
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
Guiding LLMs to Generate High-Fidelity and High-Quality Counterfactual Explanations for Text Classification
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2025)
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2025)
Anchored Alignment for Self-Explanations Enhancement
von: Villa-Arenas, Luis Felipe, et al.
Veröffentlicht: (2024)
von: Villa-Arenas, Luis Felipe, et al.
Veröffentlicht: (2024)
Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
von: Wang, Qianli, et al.
Veröffentlicht: (2026)
CoXQL: A Dataset for Parsing Explanation Requests in Conversational XAI Systems
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
von: Wang, Qianli, et al.
Veröffentlicht: (2025)
von: Wang, Qianli, et al.
Veröffentlicht: (2025)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2022)
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2022)
Multilingual Datasets for Custom Input Extraction and Explanation Requests Parsing in Conversational XAI Systems
von: Wang, Qianli, et al.
Veröffentlicht: (2025)
von: Wang, Qianli, et al.
Veröffentlicht: (2025)
LLMCheckup: Conversational Examination of Large Language Models via Interpretability Tools and Self-Explanations
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
von: Wang, Qianli, et al.
Veröffentlicht: (2024)
Judge Circuits
von: Feldhus, Nils, et al.
Veröffentlicht: (2026)
von: Feldhus, Nils, et al.
Veröffentlicht: (2026)
Proceedings of the ISCA/ITG Workshop on Diversity in Large Speech and Language Models
von: Möller, Sebastian, et al.
Veröffentlicht: (2025)
von: Möller, Sebastian, et al.
Veröffentlicht: (2025)
Enhancing Fact Retrieval in PLMs through Truthfulness
von: Youssef, Paul, et al.
Veröffentlicht: (2024)
von: Youssef, Paul, et al.
Veröffentlicht: (2024)
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
von: Yang, Jing, et al.
Veröffentlicht: (2026)
von: Yang, Jing, et al.
Veröffentlicht: (2026)
Persona Prompting as a Lens on LLM Social Reasoning
von: Yang, Jing, et al.
Veröffentlicht: (2026)
von: Yang, Jing, et al.
Veröffentlicht: (2026)
Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization
von: Sun, Jingyi, et al.
Veröffentlicht: (2026)
von: Sun, Jingyi, et al.
Veröffentlicht: (2026)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
Interpreting Language Models Through Concept Descriptions: A Survey
von: Feldhus, Nils, et al.
Veröffentlicht: (2025)
von: Feldhus, Nils, et al.
Veröffentlicht: (2025)
Simplifying Outcomes of Language Model Component Analyses with ELIA
von: Eidt, Aaron Louis, et al.
Veröffentlicht: (2026)
von: Eidt, Aaron Louis, et al.
Veröffentlicht: (2026)
Comparative Explanations: Explanation Guided Decision Making for Human-in-the-Loop Preference Selection
von: Chakraborty, Tanmay, et al.
Veröffentlicht: (2025)
von: Chakraborty, Tanmay, et al.
Veröffentlicht: (2025)
Gendered Prompting and LLM Code Review: How Gender Cues in the Prompt Shape Code Quality and Evaluation
von: Janzen, Lynn, et al.
Veröffentlicht: (2026)
von: Janzen, Lynn, et al.
Veröffentlicht: (2026)
Free-text Rationale Generation under Readability Level Control
von: Hsu, Yi-Sheng, et al.
Veröffentlicht: (2024)
von: Hsu, Yi-Sheng, et al.
Veröffentlicht: (2024)
Hybrid Annotation for Propaganda Detection: Integrating LLM Pre-Annotations with Human Intelligence
von: Sahitaj, Ariana, et al.
Veröffentlicht: (2025)
von: Sahitaj, Ariana, et al.
Veröffentlicht: (2025)
Robustness of Selected Learning Models under Label-Flipping Attack
von: Bhargava, Sarvagya, et al.
Veröffentlicht: (2025)
von: Bhargava, Sarvagya, et al.
Veröffentlicht: (2025)
Cost-Minimized Label-Flipping Poisoning Attack to LLM Alignment
von: Kusaka, Shigeki, et al.
Veröffentlicht: (2025)
von: Kusaka, Shigeki, et al.
Veröffentlicht: (2025)
Gender Bias in Explainability: Investigating Performance Disparity in Post-hoc Methods
von: Dhaini, Mahdi, et al.
Veröffentlicht: (2025)
von: Dhaini, Mahdi, et al.
Veröffentlicht: (2025)
When Truthful Representations Flip Under Deceptive Instructions?
von: Long, Xianxuan, et al.
Veröffentlicht: (2025)
von: Long, Xianxuan, et al.
Veröffentlicht: (2025)
Table Understanding and (Multimodal) LLMs: A Cross-Domain Case Study on Scientific vs. Non-Scientific Data
von: Borisova, Ekaterina, et al.
Veröffentlicht: (2025)
von: Borisova, Ekaterina, et al.
Veröffentlicht: (2025)
Persuasion Tokens for Editing Factual Knowledge in LLMs
von: Youssef, Paul, et al.
Veröffentlicht: (2026)
von: Youssef, Paul, et al.
Veröffentlicht: (2026)
The Queen of England is not England's Queen: On the Lack of Factual Coherency in PLMs
von: Youssef, Paul, et al.
Veröffentlicht: (2024)
von: Youssef, Paul, et al.
Veröffentlicht: (2024)
Explainable Bayesian Optimization
von: Chakraborty, Tanmay, et al.
Veröffentlicht: (2024)
von: Chakraborty, Tanmay, et al.
Veröffentlicht: (2024)
Towards Interpretable Deep Neural Networks for Tabular Data
von: Elhadri, Khawla, et al.
Veröffentlicht: (2025)
von: Elhadri, Khawla, et al.
Veröffentlicht: (2025)
XNNTab -- Interpretable Neural Networks for Tabular Data using Sparse Autoencoders
von: Elhadri, Khawla, et al.
Veröffentlicht: (2025)
von: Elhadri, Khawla, et al.
Veröffentlicht: (2025)
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
Finding Optimal Diverse Feature Sets with Alternative Feature Selection
von: Bach, Jakob
Veröffentlicht: (2023)
von: Bach, Jakob
Veröffentlicht: (2023)
Model-Free Counterfactual Subset Selection at Scale
von: Nguyen, Minh Hieu, et al.
Veröffentlicht: (2025)
von: Nguyen, Minh Hieu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
von: Wang, Qianli, et al.
Veröffentlicht: (2026) -
FitCF: A Framework for Automatic Feature Importance-guided Counterfactual Example Generation
von: Wang, Qianli, et al.
Veröffentlicht: (2025) -
iFlip: Iterative Feedback-driven Counterfactual Example Refinement
von: Wang, Yilong, et al.
Veröffentlicht: (2026) -
CEval: A Benchmark for Evaluating Counterfactual Text Generation
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2024) -
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2024)