AGI-Elo: How Far Are We From Mastering A Task?

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Sun, Shuo, Zhao, Yimin, Lee, Christina Dao Wen, Sun, Jiawei, Yuan, Chengran, Huang, Zefan, Li, Dongen, Yeoh, Justin KW, Prakash, Alok, Malone, Thomas W., Ang Jr, Marcelo H.
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866915303098679296
author Sun, Shuo
Zhao, Yimin
Lee, Christina Dao Wen
Sun, Jiawei
Yuan, Chengran
Huang, Zefan
Li, Dongen
Yeoh, Justin KW
Prakash, Alok
Malone, Thomas W.
Ang Jr, Marcelo H.
author_facet Sun, Shuo
Zhao, Yimin
Lee, Christina Dao Wen
Sun, Jiawei
Yuan, Chengran
Huang, Zefan
Li, Dongen
Yeoh, Justin KW
Prakash, Alok
Malone, Thomas W.
Ang Jr, Marcelo H.
contents As the field progresses toward Artificial General Intelligence (AGI), there is a pressing need for more comprehensive and insightful evaluation frameworks that go beyond aggregate performance metrics. This paper introduces a unified rating system that jointly models the difficulty of individual test cases and the competency of AI models (or humans) across vision, language, and action domains. Unlike existing metrics that focus solely on models, our approach allows for fine-grained, difficulty-aware evaluations through competitive interactions between models and tasks, capturing both the long-tail distribution of real-world challenges and the competency gap between current models and full task mastery. We validate the generalizability and robustness of our system through extensive experiments on multiple established datasets and models across distinct AGI domains. The resulting rating distributions offer novel perspectives and interpretable insights into task difficulty, model progression, and the outstanding challenges that remain on the path to achieving full AGI task mastery.
format Preprint
id arxiv_https___arxiv_org_abs_2505_12844
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AGI-Elo: How Far Are We From Mastering A Task?
Sun, Shuo
Zhao, Yimin
Lee, Christina Dao Wen
Sun, Jiawei
Yuan, Chengran
Huang, Zefan
Li, Dongen
Yeoh, Justin KW
Prakash, Alok
Malone, Thomas W.
Ang Jr, Marcelo H.
Artificial Intelligence
Robotics
As the field progresses toward Artificial General Intelligence (AGI), there is a pressing need for more comprehensive and insightful evaluation frameworks that go beyond aggregate performance metrics. This paper introduces a unified rating system that jointly models the difficulty of individual test cases and the competency of AI models (or humans) across vision, language, and action domains. Unlike existing metrics that focus solely on models, our approach allows for fine-grained, difficulty-aware evaluations through competitive interactions between models and tasks, capturing both the long-tail distribution of real-world challenges and the competency gap between current models and full task mastery. We validate the generalizability and robustness of our system through extensive experiments on multiple established datasets and models across distinct AGI domains. The resulting rating distributions offer novel perspectives and interpretable insights into task difficulty, model progression, and the outstanding challenges that remain on the path to achieving full AGI task mastery.
title AGI-Elo: How Far Are We From Mastering A Task?
topic Artificial Intelligence
Robotics
url https://arxiv.org/abs/2505.12844