Eureka: Evaluating and Understanding Large Foundation Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Balachandran, Vidhisha, Chen, Jingya, Joshi, Neel, Nushi, Besmira, Palangi, Hamid, Salinas, Eduardo, Vineet, Vibhav, Woffinden-Luey, James, Yousefi, Safoora
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913503367921664
author Balachandran, Vidhisha
Chen, Jingya
Joshi, Neel
Nushi, Besmira
Palangi, Hamid
Salinas, Eduardo
Vineet, Vibhav
Woffinden-Luey, James
Yousefi, Safoora
author_facet Balachandran, Vidhisha
Chen, Jingya
Joshi, Neel
Nushi, Besmira
Palangi, Hamid
Salinas, Eduardo
Vineet, Vibhav
Woffinden-Luey, James
Yousefi, Safoora
contents Rigorous and reproducible evaluation is critical for assessing the state of the art and for guiding scientific advances in Artificial Intelligence. Evaluation is challenging in practice due to several reasons, including benchmark saturation, lack of transparency in methods used for measurement, development challenges in extracting measurements for generative tasks, and, more generally, the extensive number of capabilities required for a well-rounded comparison across models. We make three contributions to alleviate the above challenges. First, we present Eureka, an open-source framework for standardizing evaluations of large foundation models beyond single-score reporting and rankings. Second, we introduce Eureka-Bench as an extensible collection of benchmarks testing capabilities that (i) are still challenging for state-of-the-art models and (ii) represent fundamental but overlooked language and multimodal capabilities. The inherent space for improvement in non-saturated benchmarks enables us to discover meaningful differences between models at a capability level. Third, using Eureka, we conduct an analysis of 12 state-of-the-art models, providing in-depth insights into failure understanding and model comparison, which can be leveraged to plan targeted improvements. In contrast to recent trends in reports and leaderboards showing absolute rankings and claims for one model or another to be the best, our analysis shows that there is no such best model. Different models have different strengths, but there are models that appear more often than others as best performers for some capabilities. Despite the recent improvements, current models still struggle with several fundamental capabilities including detailed image understanding, benefiting from multimodal input when available rather than fully relying on language, factuality and grounding for information retrieval, and over refusals.
format Preprint
id arxiv_https___arxiv_org_abs_2409_10566
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Eureka: Evaluating and Understanding Large Foundation Models
Balachandran, Vidhisha
Chen, Jingya
Joshi, Neel
Nushi, Besmira
Palangi, Hamid
Salinas, Eduardo
Vineet, Vibhav
Woffinden-Luey, James
Yousefi, Safoora
Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
I.2
Rigorous and reproducible evaluation is critical for assessing the state of the art and for guiding scientific advances in Artificial Intelligence. Evaluation is challenging in practice due to several reasons, including benchmark saturation, lack of transparency in methods used for measurement, development challenges in extracting measurements for generative tasks, and, more generally, the extensive number of capabilities required for a well-rounded comparison across models. We make three contributions to alleviate the above challenges. First, we present Eureka, an open-source framework for standardizing evaluations of large foundation models beyond single-score reporting and rankings. Second, we introduce Eureka-Bench as an extensible collection of benchmarks testing capabilities that (i) are still challenging for state-of-the-art models and (ii) represent fundamental but overlooked language and multimodal capabilities. The inherent space for improvement in non-saturated benchmarks enables us to discover meaningful differences between models at a capability level. Third, using Eureka, we conduct an analysis of 12 state-of-the-art models, providing in-depth insights into failure understanding and model comparison, which can be leveraged to plan targeted improvements. In contrast to recent trends in reports and leaderboards showing absolute rankings and claims for one model or another to be the best, our analysis shows that there is no such best model. Different models have different strengths, but there are models that appear more often than others as best performers for some capabilities. Despite the recent improvements, current models still struggle with several fundamental capabilities including detailed image understanding, benefiting from multimodal input when available rather than fully relying on language, factuality and grounding for information retrieval, and over refusals.
title Eureka: Evaluating and Understanding Large Foundation Models
topic Machine Learning
Artificial Intelligence
Computation and Language
Computer Vision and Pattern Recognition
I.2
url https://arxiv.org/abs/2409.10566