Learning Compact Representations of LLM Abilities via Item Response Theory

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Jianhao, Wang, Chenxu, Zhang, Gengrui, Ye, Peng, Bai, Lei, Hu, Wei, Qu, Yuzhong, Hu, Shuyue
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909818639351808
author Chen, Jianhao
Wang, Chenxu
Zhang, Gengrui
Ye, Peng
Bai, Lei
Hu, Wei
Qu, Yuzhong
Hu, Shuyue
author_facet Chen, Jianhao
Wang, Chenxu
Zhang, Gengrui
Ye, Peng
Bai, Lei
Hu, Wei
Qu, Yuzhong
Hu, Shuyue
contents Recent years have witnessed a surge in the number of large language models (LLMs), yet efficiently managing and utilizing these vast resources remains a significant challenge. In this work, we explore how to learn compact representations of LLM abilities that can facilitate downstream tasks, such as model routing and performance prediction on new benchmarks. We frame this problem as estimating the probability that a given model will correctly answer a specific query. Inspired by the item response theory (IRT) in psychometrics, we model this probability as a function of three key factors: (i) the model's multi-skill ability vector, (2) the query's discrimination vector that separates models of differing skills, and (3) the query's difficulty scalar. To learn these parameters jointly, we introduce a Mixture-of-Experts (MoE) network that couples model- and query-level embeddings. Extensive experiments demonstrate that our approach leads to state-of-the-art performance in both model routing and benchmark accuracy prediction. Moreover, analysis validates that the learned parameters encode meaningful, interpretable information about model capabilities and query characteristics.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00844
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Learning Compact Representations of LLM Abilities via Item Response Theory
Chen, Jianhao
Wang, Chenxu
Zhang, Gengrui
Ye, Peng
Bai, Lei
Hu, Wei
Qu, Yuzhong
Hu, Shuyue
Artificial Intelligence
Recent years have witnessed a surge in the number of large language models (LLMs), yet efficiently managing and utilizing these vast resources remains a significant challenge. In this work, we explore how to learn compact representations of LLM abilities that can facilitate downstream tasks, such as model routing and performance prediction on new benchmarks. We frame this problem as estimating the probability that a given model will correctly answer a specific query. Inspired by the item response theory (IRT) in psychometrics, we model this probability as a function of three key factors: (i) the model's multi-skill ability vector, (2) the query's discrimination vector that separates models of differing skills, and (3) the query's difficulty scalar. To learn these parameters jointly, we introduce a Mixture-of-Experts (MoE) network that couples model- and query-level embeddings. Extensive experiments demonstrate that our approach leads to state-of-the-art performance in both model routing and benchmark accuracy prediction. Moreover, analysis validates that the learned parameters encode meaningful, interpretable information about model capabilities and query characteristics.
title Learning Compact Representations of LLM Abilities via Item Response Theory
topic Artificial Intelligence
url https://arxiv.org/abs/2510.00844