IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Maimon, Aviya, Cohen, Amir DN, Vishne, Gal, Ravfogel, Shauli, Tsarfaty, Reut
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911079260487680
author Maimon, Aviya
Cohen, Amir DN
Vishne, Gal
Ravfogel, Shauli
Tsarfaty, Reut
author_facet Maimon, Aviya
Cohen, Amir DN
Vishne, Gal
Ravfogel, Shauli
Tsarfaty, Reut
contents Current evaluations of large language models (LLMs) rely on benchmark scores, but it is difficult to interpret what these individual scores reveal about a model's overall skills. Specifically, as a community we lack understanding of how tasks relate to one another, what they measure in common, how they differ, or which ones are redundant. As a result, models are often assessed via a single score averaged across benchmarks, an approach that fails to capture the models' wholistic strengths and limitations. Here, we propose a new evaluation paradigm that uses factor analysis to identify latent skills driving performance across benchmarks. We apply this method to a comprehensive new leaderboard showcasing the performance of 60 LLMs on 44 tasks, and identify a small set of latent skills that largely explain performance. Finally, we turn these insights into practical tools that identify redundant tasks, aid in model selection, and profile models along each latent skill.
format Preprint
id arxiv_https___arxiv_org_abs_2507_20208
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
Maimon, Aviya
Cohen, Amir DN
Vishne, Gal
Ravfogel, Shauli
Tsarfaty, Reut
Computation and Language
Current evaluations of large language models (LLMs) rely on benchmark scores, but it is difficult to interpret what these individual scores reveal about a model's overall skills. Specifically, as a community we lack understanding of how tasks relate to one another, what they measure in common, how they differ, or which ones are redundant. As a result, models are often assessed via a single score averaged across benchmarks, an approach that fails to capture the models' wholistic strengths and limitations. Here, we propose a new evaluation paradigm that uses factor analysis to identify latent skills driving performance across benchmarks. We apply this method to a comprehensive new leaderboard showcasing the performance of 60 LLMs on 44 tasks, and identify a small set of latent skills that largely explain performance. Finally, we turn these insights into practical tools that identify redundant tasks, aid in model selection, and profile models along each latent skill.
title IQ Test for LLMs: An Evaluation Framework for Uncovering Core Skills in LLMs
topic Computation and Language
url https://arxiv.org/abs/2507.20208