Law of the Weakest Link: Cross Capabilities of Large Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhong, Ming, Zhang, Aston, Wang, Xuewei, Hou, Rui, Xiong, Wenhan, Zhu, Chenguang, Chen, Zhengxing, Tan, Liang, Bi, Chloe, Lewis, Mike, Popuri, Sravya, Narang, Sharan, Kambadur, Melanie, Mahajan, Dhruv, Edunov, Sergey, Han, Jiawei, van der Maaten, Laurens
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912056391761920
author Zhong, Ming
Zhang, Aston
Wang, Xuewei
Hou, Rui
Xiong, Wenhan
Zhu, Chenguang
Chen, Zhengxing
Tan, Liang
Bi, Chloe
Lewis, Mike
Popuri, Sravya
Narang, Sharan
Kambadur, Melanie
Mahajan, Dhruv
Edunov, Sergey
Han, Jiawei
van der Maaten, Laurens
author_facet Zhong, Ming
Zhang, Aston
Wang, Xuewei
Hou, Rui
Xiong, Wenhan
Zhu, Chenguang
Chen, Zhengxing
Tan, Liang
Bi, Chloe
Lewis, Mike
Popuri, Sravya
Narang, Sharan
Kambadur, Melanie
Mahajan, Dhruv
Edunov, Sergey
Han, Jiawei
van der Maaten, Laurens
contents The development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term cross capabilities. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce CrossEval, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that, in both static evaluations and attempts to enhance specific abilities, current LLMs consistently exhibit the "Law of the Weakest Link," where cross-capability performance is significantly constrained by the weakest component. Specifically, across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight the under-performance of LLMs in cross-capability tasks, making the identification and improvement of the weakest capabilities a critical priority for future research to optimize performance in complex, multi-dimensional scenarios.
format Preprint
id arxiv_https___arxiv_org_abs_2409_19951
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Law of the Weakest Link: Cross Capabilities of Large Language Models
Zhong, Ming
Zhang, Aston
Wang, Xuewei
Hou, Rui
Xiong, Wenhan
Zhu, Chenguang
Chen, Zhengxing
Tan, Liang
Bi, Chloe
Lewis, Mike
Popuri, Sravya
Narang, Sharan
Kambadur, Melanie
Mahajan, Dhruv
Edunov, Sergey
Han, Jiawei
van der Maaten, Laurens
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
The development and evaluation of Large Language Models (LLMs) have largely focused on individual capabilities. However, this overlooks the intersection of multiple abilities across different types of expertise that are often required for real-world tasks, which we term cross capabilities. To systematically explore this concept, we first define seven core individual capabilities and then pair them to form seven common cross capabilities, each supported by a manually constructed taxonomy. Building on these definitions, we introduce CrossEval, a benchmark comprising 1,400 human-annotated prompts, with 100 prompts for each individual and cross capability. To ensure reliable evaluation, we involve expert annotators to assess 4,200 model responses, gathering 8,400 human ratings with detailed explanations to serve as reference examples. Our findings reveal that, in both static evaluations and attempts to enhance specific abilities, current LLMs consistently exhibit the "Law of the Weakest Link," where cross-capability performance is significantly constrained by the weakest component. Specifically, across 58 cross-capability scores from 17 models, 38 scores are lower than all individual capabilities, while 20 fall between strong and weak, but closer to the weaker ability. These results highlight the under-performance of LLMs in cross-capability tasks, making the identification and improvement of the weakest capabilities a critical priority for future research to optimize performance in complex, multi-dimensional scenarios.
title Law of the Weakest Link: Cross Capabilities of Large Language Models
topic Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2409.19951