LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
Fuente:
arXiv
Salvato in:
| Autori principali: | Toker, Gilat, Calderon, Nitay, Amosy, Ohad, Reichart, Roi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs
di: Calderon, Nitay, et al.
Pubblicazione: (2024)
di: Calderon, Nitay, et al.
Pubblicazione: (2024)
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
di: Calderon, Nitay, et al.
Pubblicazione: (2025)
di: Calderon, Nitay, et al.
Pubblicazione: (2025)
NL-Eye: Abductive NLI for Images
di: Ventura, Mor, et al.
Pubblicazione: (2024)
di: Ventura, Mor, et al.
Pubblicazione: (2024)
The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth
di: Lissak, Shir, et al.
Pubblicazione: (2024)
di: Lissak, Shir, et al.
Pubblicazione: (2024)
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
di: Ventura, Mor, et al.
Pubblicazione: (2025)
di: Ventura, Mor, et al.
Pubblicazione: (2025)
Multi-Domain Explainability of Preferences
di: Calderon, Nitay, et al.
Pubblicazione: (2025)
di: Calderon, Nitay, et al.
Pubblicazione: (2025)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
Are LLMs Better than Reported? Detecting Label Errors and Mitigating Their Effect on Model Performance
di: Nahum, Omer, et al.
Pubblicazione: (2024)
di: Nahum, Omer, et al.
Pubblicazione: (2024)
Systematic Biases in LLM Simulations of Debates
di: Taubenfeld, Amir, et al.
Pubblicazione: (2024)
di: Taubenfeld, Amir, et al.
Pubblicazione: (2024)
Few-Shot Knowledge Distillation of LLMs With Counterfactual Explanations
di: Hamman, Faisal, et al.
Pubblicazione: (2025)
di: Hamman, Faisal, et al.
Pubblicazione: (2025)
GLEE: A Unified Framework and Benchmark for Language-based Economic Environments
di: Shapira, Eilam, et al.
Pubblicazione: (2024)
di: Shapira, Eilam, et al.
Pubblicazione: (2024)
Predicting Decisions of AI Agents from Limited Interaction through Text-Tabular Modeling
di: Shapira, Eilam, et al.
Pubblicazione: (2026)
di: Shapira, Eilam, et al.
Pubblicazione: (2026)
Counterfactual Simulatability of LLM Explanations for Generation Tasks
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
di: Limpijankit, Marvin, et al.
Pubblicazione: (2025)
Can LLMs Replace Economic Choice Prediction Labs? The Case of Language-based Persuasion Games
di: Shapira, Eilam, et al.
Pubblicazione: (2024)
di: Shapira, Eilam, et al.
Pubblicazione: (2024)
The Poisoned Apple Effect: Strategic Manipulation of Mediated Markets via Technology Expansion of AI Agents
di: Shapira, Eilam, et al.
Pubblicazione: (2026)
di: Shapira, Eilam, et al.
Pubblicazione: (2026)
AdaptiVocab: Enhancing LLM Efficiency in Focused Domains through Lightweight Vocabulary Adaptation
di: Nakash, Itay, et al.
Pubblicazione: (2025)
di: Nakash, Itay, et al.
Pubblicazione: (2025)
Benchmarking Concept-Spilling Across Languages in LLMs
di: Badanin, Ilia, et al.
Pubblicazione: (2026)
di: Badanin, Ilia, et al.
Pubblicazione: (2026)
A Comparative Analysis of Counterfactual Explanation Methods for Text Classifiers
di: McAleese, Stephen, et al.
Pubblicazione: (2024)
di: McAleese, Stephen, et al.
Pubblicazione: (2024)
Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality
di: Calderon, Nitay, et al.
Pubblicazione: (2026)
di: Calderon, Nitay, et al.
Pubblicazione: (2026)
Navigating Cultural Chasms: Exploring and Unlocking the Cultural POV of Text-To-Image Models
di: Ventura, Mor, et al.
Pubblicazione: (2023)
di: Ventura, Mor, et al.
Pubblicazione: (2023)
LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
di: Mayne, Harry, et al.
Pubblicazione: (2025)
di: Mayne, Harry, et al.
Pubblicazione: (2025)
Donors and Recipients: On Asymmetric Transfer Across Tasks and Languages with Parameter-Efficient Fine-Tuning
di: Dymkiewicz, Kajetan, et al.
Pubblicazione: (2025)
di: Dymkiewicz, Kajetan, et al.
Pubblicazione: (2025)
MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
di: Chen, Wenting, et al.
Pubblicazione: (2026)
di: Chen, Wenting, et al.
Pubblicazione: (2026)
Dementia Through Different Eyes: Explainable Modeling of Human and LLM Perceptions for Early Awareness
di: Peled-Cohen, Lotem, et al.
Pubblicazione: (2025)
di: Peled-Cohen, Lotem, et al.
Pubblicazione: (2025)
NoisyCausal: A Benchmark for Evaluating Causal Reasoning Under Structured Noise
di: Xu, Zhi, et al.
Pubblicazione: (2026)
di: Xu, Zhi, et al.
Pubblicazione: (2026)
Pretrained LLMs Learn Multiple Types of Uncertainty
di: Cohen, Roi, et al.
Pubblicazione: (2025)
di: Cohen, Roi, et al.
Pubblicazione: (2025)
Natural Language Counterfactual Explanations for Graphs Using Large Language Models
di: Giorgi, Flavio, et al.
Pubblicazione: (2024)
di: Giorgi, Flavio, et al.
Pubblicazione: (2024)
Can LLMs Explain Themselves Counterfactually?
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2025)
di: Dehghanighobadi, Zahra, et al.
Pubblicazione: (2025)
Generative Framework for Personalized Persuasion: Inferring Causal, Counterfactual, and Latent Knowledge
di: Zeng, Donghuo, et al.
Pubblicazione: (2025)
di: Zeng, Donghuo, et al.
Pubblicazione: (2025)
RATE: Causal Explainability of Reward Models with Imperfect Counterfactuals
di: Reber, David, et al.
Pubblicazione: (2024)
di: Reber, David, et al.
Pubblicazione: (2024)
LLMs for Generating and Evaluating Counterfactuals: A Comprehensive Study
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024)
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024)
CEval: A Benchmark for Evaluating Counterfactual Text Generation
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024)
di: Nguyen, Van Bach, et al.
Pubblicazione: (2024)
Evaluating Causal Explanation in Medical Reports with LLM-Based and Human-Aligned Metrics
di: Cho, Yousang, et al.
Pubblicazione: (2025)
di: Cho, Yousang, et al.
Pubblicazione: (2025)
Benchmarking LLMs for Pairwise Causal Discovery in Biomedical and Multi-Domain Contexts
di: Anuyah, Sydney, et al.
Pubblicazione: (2026)
di: Anuyah, Sydney, et al.
Pubblicazione: (2026)
Beyond Surface Structure: A Causal Assessment of LLMs' Comprehension Ability
di: Han, Yujin, et al.
Pubblicazione: (2024)
di: Han, Yujin, et al.
Pubblicazione: (2024)
Leveraging Prompt-Learning for Structured Information Extraction from Crohn's Disease Radiology Reports in a Low-Resource Language
di: Hazan, Liam, et al.
Pubblicazione: (2024)
di: Hazan, Liam, et al.
Pubblicazione: (2024)
JAILJUDGE: A Comprehensive Jailbreak Judge Benchmark with Multi-Agent Enhanced Explanation Evaluation Framework
di: Liu, Fan, et al.
Pubblicazione: (2024)
di: Liu, Fan, et al.
Pubblicazione: (2024)
Local Explanations and Self-Explanations for Assessing Faithfulness in black-box LLMs
di: Fragkathoulas, Christos, et al.
Pubblicazione: (2024)
di: Fragkathoulas, Christos, et al.
Pubblicazione: (2024)
Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric Assessments
di: Domnich, Marharyta, et al.
Pubblicazione: (2024)
di: Domnich, Marharyta, et al.
Pubblicazione: (2024)
Prompt-Counterfactual Explanations for Generative AI System Behavior
di: Goethals, Sofie, et al.
Pubblicazione: (2026)
di: Goethals, Sofie, et al.
Pubblicazione: (2026)
Documenti analoghi
-
On Behalf of the Stakeholders: Trends in NLP Model Interpretability in the Era of LLMs
di: Calderon, Nitay, et al.
Pubblicazione: (2024) -
The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
di: Calderon, Nitay, et al.
Pubblicazione: (2025) -
NL-Eye: Abductive NLI for Images
di: Ventura, Mor, et al.
Pubblicazione: (2024) -
The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth
di: Lissak, Shir, et al.
Pubblicazione: (2024) -
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
di: Ventura, Mor, et al.
Pubblicazione: (2025)