Explainable Benchmarking through the Lense of Concept Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Quannian, Röder, Michael, Srivastava, Nikit, Kouagou, N'Dah Jean, Ngomo, Axel-Cyrille Ngonga
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908607072698368
author Zhang, Quannian
Röder, Michael
Srivastava, Nikit
Kouagou, N'Dah Jean
Ngomo, Axel-Cyrille Ngonga
author_facet Zhang, Quannian
Röder, Michael
Srivastava, Nikit
Kouagou, N'Dah Jean
Ngomo, Axel-Cyrille Ngonga
contents Evaluating competing systems in a comparable way, i.e., benchmarking them, is an undeniable pillar of the scientific method. However, system performance is often summarized via a small number of metrics. The analysis of the evaluation details and the derivation of insights for further development or use remains a tedious manual task with often biased results. Thus, this paper argues for a new type of benchmarking, which is dubbed explainable benchmarking. The aim of explainable benchmarking approaches is to automatically generate explanations for the performance of systems in a benchmark. We provide a first instantiation of this paradigm for knowledge-graph-based question answering systems. We compute explanations by using a novel concept learning approach developed for large knowledge graphs called PruneCEL. Our evaluation shows that PruneCEL outperforms state-of-the-art concept learners on the task of explainable benchmarking by up to 0.55 points F1 measure. A task-driven user study with 41 participants shows that in 80\% of the cases, the majority of participants can accurately predict the behavior of a system based on our explanations. Our code and data are available at https://github.com/dice-group/PruneCEL/tree/K-cap2025
format Preprint
id arxiv_https___arxiv_org_abs_2510_20439
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Explainable Benchmarking through the Lense of Concept Learning
Zhang, Quannian
Röder, Michael
Srivastava, Nikit
Kouagou, N'Dah Jean
Ngomo, Axel-Cyrille Ngonga
Machine Learning
Evaluating competing systems in a comparable way, i.e., benchmarking them, is an undeniable pillar of the scientific method. However, system performance is often summarized via a small number of metrics. The analysis of the evaluation details and the derivation of insights for further development or use remains a tedious manual task with often biased results. Thus, this paper argues for a new type of benchmarking, which is dubbed explainable benchmarking. The aim of explainable benchmarking approaches is to automatically generate explanations for the performance of systems in a benchmark. We provide a first instantiation of this paradigm for knowledge-graph-based question answering systems. We compute explanations by using a novel concept learning approach developed for large knowledge graphs called PruneCEL. Our evaluation shows that PruneCEL outperforms state-of-the-art concept learners on the task of explainable benchmarking by up to 0.55 points F1 measure. A task-driven user study with 41 participants shows that in 80\% of the cases, the majority of participants can accurately predict the behavior of a system based on our explanations. Our code and data are available at https://github.com/dice-group/PruneCEL/tree/K-cap2025
title Explainable Benchmarking through the Lense of Concept Learning
topic Machine Learning
url https://arxiv.org/abs/2510.20439