The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Meimandi, Kiana Jafari, Aránguiz-Dias, Gabriela, Kim, Grace Ra, Saadeddin, Lana, Griffith, Allie, Kochenderfer, Mykel J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912621916061696
author Meimandi, Kiana Jafari
Aránguiz-Dias, Gabriela
Kim, Grace Ra
Saadeddin, Lana
Griffith, Allie
Kochenderfer, Mykel J.
author_facet Meimandi, Kiana Jafari
Aránguiz-Dias, Gabriela
Kim, Grace Ra
Saadeddin, Lana
Griffith, Allie
Kochenderfer, Mykel J.
contents As industry reports claim agentic AI systems deliver double-digit productivity gains and multi-trillion dollar economic potential, the validity of these claims has become critical for investment decisions, regulatory policy, and responsible technology adoption. However, this paper demonstrates that current evaluation practices for agentic AI systems exhibit a systemic imbalance that calls into question prevailing industry productivity claims. Our systematic review of 84 papers (2023--2025) reveals an evaluation imbalance where technical metrics dominate assessments (83%), while human-centered (30%), safety (53%), and economic assessments (30%) remain peripheral, with only 15% incorporating both technical and human dimensions. This measurement gap creates a fundamental disconnect between benchmark success and deployment value. We present evidence from healthcare, finance, and retail sectors where systems excelling on technical metrics failed in real-world implementation due to unmeasured human, temporal, and contextual factors. Our position is not against agentic AI's potential, but rather that current evaluation frameworks systematically privilege narrow technical metrics while neglecting dimensions critical to real-world success. We propose a balanced four-axis evaluation model and call on the community to lead this paradigm shift because benchmark-driven optimization shapes what we build. By redefining evaluation practices, we can better align industry claims with deployment realities and ensure responsible scaling of agentic systems in high-stakes domains.
format Preprint
id arxiv_https___arxiv_org_abs_2506_02064
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
Meimandi, Kiana Jafari
Aránguiz-Dias, Gabriela
Kim, Grace Ra
Saadeddin, Lana
Griffith, Allie
Kochenderfer, Mykel J.
Computers and Society
Human-Computer Interaction
As industry reports claim agentic AI systems deliver double-digit productivity gains and multi-trillion dollar economic potential, the validity of these claims has become critical for investment decisions, regulatory policy, and responsible technology adoption. However, this paper demonstrates that current evaluation practices for agentic AI systems exhibit a systemic imbalance that calls into question prevailing industry productivity claims. Our systematic review of 84 papers (2023--2025) reveals an evaluation imbalance where technical metrics dominate assessments (83%), while human-centered (30%), safety (53%), and economic assessments (30%) remain peripheral, with only 15% incorporating both technical and human dimensions. This measurement gap creates a fundamental disconnect between benchmark success and deployment value. We present evidence from healthcare, finance, and retail sectors where systems excelling on technical metrics failed in real-world implementation due to unmeasured human, temporal, and contextual factors. Our position is not against agentic AI's potential, but rather that current evaluation frameworks systematically privilege narrow technical metrics while neglecting dimensions critical to real-world success. We propose a balanced four-axis evaluation model and call on the community to lead this paradigm shift because benchmark-driven optimization shapes what we build. By redefining evaluation practices, we can better align industry claims with deployment realities and ensure responsible scaling of agentic systems in high-stakes domains.
title The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
topic Computers and Society
Human-Computer Interaction
url https://arxiv.org/abs/2506.02064