AGI-Elo: How Far Are We From Mastering A Task?
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Shuo, Zhao, Yimin, Lee, Christina Dao Wen, Sun, Jiawei, Yuan, Chengran, Huang, Zefan, Li, Dongen, Yeoh, Justin KW, Prakash, Alok, Malone, Thomas W., Ang Jr, Marcelo H. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PIE: Perception and Interaction Enhanced End-to-End Motion Planning for Autonomous Driving
by: Yuan, Chengran, et al.
Published: (2025)
by: Yuan, Chengran, et al.
Published: (2025)
DRAMA: An Efficient End-to-end Motion Planner for Autonomous Driving with Mamba
by: Yuan, Chengran, et al.
Published: (2024)
by: Yuan, Chengran, et al.
Published: (2024)
CARLA-Loc: Synthetic SLAM Dataset with Full-stack Sensor Setup in Challenging Weather and Dynamic Environments
by: Han, Yuhang, et al.
Published: (2023)
by: Han, Yuhang, et al.
Published: (2023)
DriveSceneGen: Generating Diverse and Realistic Driving Scenarios from Scratch
by: Sun, Shuo, et al.
Published: (2023)
by: Sun, Shuo, et al.
Published: (2023)
RMP-YOLO: A Robust Motion Predictor for Partially Observable Scenarios even if You Only Look Once
by: Sun, Jiawei, et al.
Published: (2024)
by: Sun, Jiawei, et al.
Published: (2024)
DINO-CoDT: Multi-class Collaborative Detection and Tracking with Vision Foundation Models
by: He, Xunjie, et al.
Published: (2025)
by: He, Xunjie, et al.
Published: (2025)
How Far Are We From AGI: Are LLMs All We Need?
by: Feng, Tao, et al.
Published: (2024)
by: Feng, Tao, et al.
Published: (2024)
RMGS-SLAM: Real-time Multi-sensor Gaussian Splatting SLAM
by: Li, Dongen, et al.
Published: (2026)
by: Li, Dongen, et al.
Published: (2026)
ControlMTR: Control-Guided Motion Transformer with Scene-Compliant Intention Points for Feasible Motion Prediction
by: Sun, Jiawei, et al.
Published: (2024)
by: Sun, Jiawei, et al.
Published: (2024)
ADM: Accelerated Diffusion Model via Estimated Priors for Robust Motion Prediction under Uncertainties
by: Li, Jiahui, et al.
Published: (2024)
by: Li, Jiahui, et al.
Published: (2024)
IMPACT: Behavioral Intention-aware Multimodal Trajectory Prediction with Adaptive Context Trimming
by: Sun, Jiawei, et al.
Published: (2025)
by: Sun, Jiawei, et al.
Published: (2025)
Automated Testing of Task-based Chatbots: How Far Are We?
by: Clerissi, Diego, et al.
Published: (2026)
by: Clerissi, Diego, et al.
Published: (2026)
Vulnerability-Affected Versions Identification: How Far Are We?
by: Chen, Xingchu, et al.
Published: (2025)
by: Chen, Xingchu, et al.
Published: (2025)
The Digital Cybersecurity Expert: How Far Have We Come?
by: Wang, Dawei, et al.
Published: (2025)
by: Wang, Dawei, et al.
Published: (2025)
An Empirical Study on Automatically Detecting AI-Generated Source Code: How Far Are We?
by: Suh, Hyunjae, et al.
Published: (2024)
by: Suh, Hyunjae, et al.
Published: (2024)
PICABench: How Far Are We from Physically Realistic Image Editing?
by: Pu, Yuandong, et al.
Published: (2025)
by: Pu, Yuandong, et al.
Published: (2025)
How Far Are We from True Unlearnability?
by: Ye, Kai, et al.
Published: (2025)
by: Ye, Kai, et al.
Published: (2025)
LLMs for Relational Reasoning: How Far are We?
by: Li, Zhiming, et al.
Published: (2024)
by: Li, Zhiming, et al.
Published: (2024)
Effects of Dietary Fermented Garlic on the Growth Performance, Relative Organ Weights, Intestinal Morphology, Cecal Microflora and Serum Characteristics of Broiler Chickens
by: KW Lee
Published: (2016)
by: KW Lee
Published: (2016)
Static Application Security Testing (SAST) Tools for Smart Contracts: How Far Are We?
by: Li, Kaixuan, et al.
Published: (2024)
by: Li, Kaixuan, et al.
Published: (2024)
How Far Are We From True Auto-Research?
by: Zhang, Zhengxin, et al.
Published: (2026)
by: Zhang, Zhengxin, et al.
Published: (2026)
How Far Are We from Optimal Reasoning Efficiency?
by: Gao, Jiaxuan, et al.
Published: (2025)
by: Gao, Jiaxuan, et al.
Published: (2025)
Duplicate Bug Report Detection: How Far Are We?
by: Zhang, Ting, et al.
Published: (2022)
by: Zhang, Ting, et al.
Published: (2022)
Retrieval-Augmented Test Generation: How Far Are We?
by: Shin, Jiho, et al.
Published: (2024)
by: Shin, Jiho, et al.
Published: (2024)
Portfolio Assessment: How Far Have We Come?
by: Brown, Carol A.
Published: (2002)
by: Brown, Carol A.
Published: (2002)
Perception-Aware Multimodal Spatial Reasoning from Monocular Images
by: Cheng, Yanchun, et al.
Published: (2026)
by: Cheng, Yanchun, et al.
Published: (2026)
Can AI Agents Generate Microservices? How Far are We?
by: Adnan, Bassam, et al.
Published: (2026)
by: Adnan, Bassam, et al.
Published: (2026)
Vulnerability Detection with Code Language Models: How Far Are We?
by: Ding, Yangruibo, et al.
Published: (2024)
by: Ding, Yangruibo, et al.
Published: (2024)
Representation Learning for Stack Overflow Posts: How Far are We?
by: He, Junda, et al.
Published: (2023)
by: He, Junda, et al.
Published: (2023)
How Far Are We from Intelligent Visual Deductive Reasoning?
by: Zhang, Yizhe, et al.
Published: (2024)
by: Zhang, Yizhe, et al.
Published: (2024)
Model Editing for LLMs4Code: How Far are We?
by: Li, Xiaopeng, et al.
Published: (2024)
by: Li, Xiaopeng, et al.
Published: (2024)
Using LLMs for Security Advisory Investigations: How Far Are We?
by: Abdullah, Bayu Fedra, et al.
Published: (2025)
by: Abdullah, Bayu Fedra, et al.
Published: (2025)
The AI Hippocampus: How Far are We From Human Memory?
by: Jia, Zixia, et al.
Published: (2026)
by: Jia, Zixia, et al.
Published: (2026)
LLM For Loop Invariant Generation and Fixing: How Far Are We?
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
by: Akhond, Mostafijur Rahman, et al.
Published: (2025)
Fairness Improvement with Multiple Protected Attributes: How Far Are We?
by: Chen, Zhenpeng, et al.
Published: (2023)
by: Chen, Zhenpeng, et al.
Published: (2023)
AGI, Governments, and Free Societies
by: Bullock, Justin B., et al.
Published: (2025)
by: Bullock, Justin B., et al.
Published: (2025)
Keep the Future Human: Why and How We Should Close the Gates to AGI and Superintelligence, and What We Should Build Instead
by: Aguirre, Anthony
Published: (2023)
by: Aguirre, Anthony
Published: (2023)
Status and assessment of humphead wrase, Cheilinus undulates stock in Kenya.
by: Nyaga, K.W.
Published: (2009)
by: Nyaga, K.W.
Published: (2009)
Large Language Models for Equivalent Mutant Detection: How Far Are We?
by: Tian, Zhao, et al.
Published: (2024)
by: Tian, Zhao, et al.
Published: (2024)
Measuring Emergent Capabilities of LLMs for Software Engineering: How Far Are We?
by: O'Brien, Conor, et al.
Published: (2024)
by: O'Brien, Conor, et al.
Published: (2024)
Similar Items
-
PIE: Perception and Interaction Enhanced End-to-End Motion Planning for Autonomous Driving
by: Yuan, Chengran, et al.
Published: (2025) -
DRAMA: An Efficient End-to-end Motion Planner for Autonomous Driving with Mamba
by: Yuan, Chengran, et al.
Published: (2024) -
CARLA-Loc: Synthetic SLAM Dataset with Full-stack Sensor Setup in Challenging Weather and Dynamic Environments
by: Han, Yuhang, et al.
Published: (2023) -
DriveSceneGen: Generating Diverse and Realistic Driving Scenarios from Scratch
by: Sun, Shuo, et al.
Published: (2023) -
RMP-YOLO: A Robust Motion Predictor for Partially Observable Scenarios even if You Only Look Once
by: Sun, Jiawei, et al.
Published: (2024)