EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dhaini, Mahdi, Hussain, Kafaite Zahra, Zaradoukas, Efstratios, Kasneci, Gjergji
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916717362413568
author Dhaini, Mahdi
Hussain, Kafaite Zahra
Zaradoukas, Efstratios
Kasneci, Gjergji
author_facet Dhaini, Mahdi
Hussain, Kafaite Zahra
Zaradoukas, Efstratios
Kasneci, Gjergji
contents As Natural Language Processing (NLP) models continue to evolve and become integral to high-stakes applications, ensuring their interpretability remains a critical challenge. Given the growing variety of explainability methods and diverse stakeholder requirements, frameworks that help stakeholders select appropriate explanations tailored to their specific use cases are increasingly important. To address this need, we introduce EvalxNLP, a Python framework for benchmarking state-of-the-art feature attribution methods for transformer-based NLP models. EvalxNLP integrates eight widely recognized explainability techniques from the Explainable AI (XAI) literature, enabling users to generate and evaluate explanations based on key properties such as faithfulness, plausibility, and complexity. Our framework also provides interactive, LLM-based textual explanations, facilitating user understanding of the generated explanations and evaluation outcomes. Human evaluation results indicate high user satisfaction with EvalxNLP, suggesting it is a promising framework for benchmarking explanation methods across diverse user groups. By offering a user-friendly and extensible platform, EvalxNLP aims at democratizing explainability tools and supporting the systematic comparison and advancement of XAI techniques in NLP.
format Preprint
id arxiv_https___arxiv_org_abs_2505_01238
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
Dhaini, Mahdi
Hussain, Kafaite Zahra
Zaradoukas, Efstratios
Kasneci, Gjergji
Computation and Language
Artificial Intelligence
Machine Learning
As Natural Language Processing (NLP) models continue to evolve and become integral to high-stakes applications, ensuring their interpretability remains a critical challenge. Given the growing variety of explainability methods and diverse stakeholder requirements, frameworks that help stakeholders select appropriate explanations tailored to their specific use cases are increasingly important. To address this need, we introduce EvalxNLP, a Python framework for benchmarking state-of-the-art feature attribution methods for transformer-based NLP models. EvalxNLP integrates eight widely recognized explainability techniques from the Explainable AI (XAI) literature, enabling users to generate and evaluate explanations based on key properties such as faithfulness, plausibility, and complexity. Our framework also provides interactive, LLM-based textual explanations, facilitating user understanding of the generated explanations and evaluation outcomes. Human evaluation results indicate high user satisfaction with EvalxNLP, suggesting it is a promising framework for benchmarking explanation methods across diverse user groups. By offering a user-friendly and extensible platform, EvalxNLP aims at democratizing explainability tools and supporting the systematic comparison and advancement of XAI techniques in NLP.
title EvalxNLP: A Framework for Benchmarking Post-Hoc Explainability Methods on NLP Models
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2505.01238