Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Xu, Gong, Xudong, Qin, Jiacheng, Wang, Qiang, Liao, JiaQi, Wang, Zhe, Feng, Dawei, Ding, Bo
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908961819590656
author Zhang, Xu
Gong, Xudong
Qin, Jiacheng
Wang, Qiang
Liao, JiaQi
Wang, Zhe
Feng, Dawei
Ding, Bo
author_facet Zhang, Xu
Gong, Xudong
Qin, Jiacheng
Wang, Qiang
Liao, JiaQi
Wang, Zhe
Feng, Dawei
Ding, Bo
contents Current evaluations of large language models aggregate performance across diverse tasks into single scores. This obscures fine-grained ability variation, limiting targeted model improvement and ability-guided selection for specific tasks. Motivated by this gap, we propose a cognitive diagnostic framework that estimates model abilities across multiple fine-grained dimensions. For mathematics, we construct a 35-dimensional ability taxonomy grounded in cognitive theory and domain knowledge. The framework employs multidimensional Item Response Theory with an item-ability association matrix to estimate fine-grained ability levels, which in turn enable prediction of performance on unseen items (questions of benchmark). Evaluated on 41 models, our approach demonstrates strong criterion validity, consistent ability estimates across benchmarks, and accurate prediction of unseen items with AUC ranging from 0.80 to 0.89 within benchmarks and from 0.77 to 0.86 across benchmarks, substantially exceeding trivial baselines. The framework generalizes across scientific domains, producing consistent diagnostic performance in physics (27 dimensions), chemistry (58 dimensions), and computer science (12 dimensions). This work establishes a principled framework for fine-grained assessment of abilities, with potential applications in targeted training, ability-guided model selection, and ability-aware benchmark design.
format Preprint
id arxiv_https___arxiv_org_abs_2604_12191
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities
Zhang, Xu
Gong, Xudong
Qin, Jiacheng
Wang, Qiang
Liao, JiaQi
Wang, Zhe
Feng, Dawei
Ding, Bo
Artificial Intelligence
Current evaluations of large language models aggregate performance across diverse tasks into single scores. This obscures fine-grained ability variation, limiting targeted model improvement and ability-guided selection for specific tasks. Motivated by this gap, we propose a cognitive diagnostic framework that estimates model abilities across multiple fine-grained dimensions. For mathematics, we construct a 35-dimensional ability taxonomy grounded in cognitive theory and domain knowledge. The framework employs multidimensional Item Response Theory with an item-ability association matrix to estimate fine-grained ability levels, which in turn enable prediction of performance on unseen items (questions of benchmark). Evaluated on 41 models, our approach demonstrates strong criterion validity, consistent ability estimates across benchmarks, and accurate prediction of unseen items with AUC ranging from 0.80 to 0.89 within benchmarks and from 0.77 to 0.86 across benchmarks, substantially exceeding trivial baselines. The framework generalizes across scientific domains, producing consistent diagnostic performance in physics (27 dimensions), chemistry (58 dimensions), and computer science (12 dimensions). This work establishes a principled framework for fine-grained assessment of abilities, with potential applications in targeted training, ability-guided model selection, and ability-aware benchmark design.
title Beyond Scores: Diagnostic LLM Evaluation via Fine-Grained Abilities
topic Artificial Intelligence
url https://arxiv.org/abs/2604.12191