The Measurement Imbalance in Agentic AI Evaluation Undermines Industry Productivity Claims
Fuente:
arXiv
Saved in:
| Main Authors: | Meimandi, Kiana Jafari, Aránguiz-Dias, Gabriela, Kim, Grace Ra, Saadeddin, Lana, Griffith, Allie, Kochenderfer, Mykel J. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare
by: Dias, Gabriela Aránguiz, et al.
Published: (2026)
by: Dias, Gabriela Aránguiz, et al.
Published: (2026)
An Adaptive Responsible AI Governance Framework for Decentralized Organizations
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
by: Meimandi, Kiana Jafari, et al.
Published: (2025)
Responsible AI in the Global Context: Maturity Model and Survey
by: Reuel, Anka, et al.
Published: (2024)
by: Reuel, Anka, et al.
Published: (2024)
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
by: Jafari, Kiana, et al.
Published: (2026)
by: Jafari, Kiana, et al.
Published: (2026)
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
by: Hardy, Amelia, et al.
Published: (2024)
by: Hardy, Amelia, et al.
Published: (2024)
Digital Simulations to Enhance Military Medical Evacuation Decision-Making
by: Fischer, Jeremy, et al.
Published: (2025)
by: Fischer, Jeremy, et al.
Published: (2025)
Optimized but Unowned: How AI-Authored Goals Undermine the Motivation They Are Meant to Drive
by: Chi, Vivienne Bihe, et al.
Published: (2026)
by: Chi, Vivienne Bihe, et al.
Published: (2026)
AI Personalization Paradox: Personalized AI Increases Superficial Engagement in Reading while Undermines Autonomy and Ownership in Writing
by: Qin, Peinuan, et al.
Published: (2026)
by: Qin, Peinuan, et al.
Published: (2026)
LeRAAT: LLM-Enabled Real-Time Aviation Advisory Tool
by: Schlichting, Marc R., et al.
Published: (2025)
by: Schlichting, Marc R., et al.
Published: (2025)
AI for Abolition? A Participatory Design Approach
by: Wang, Carolyn, et al.
Published: (2025)
by: Wang, Carolyn, et al.
Published: (2025)
LLM Agent Meets Agentic AI: Can LLM Agents Simulate Customers to Evaluate Agentic-AI-based Shopping Assistants?
by: Sun, Lu, et al.
Published: (2025)
by: Sun, Lu, et al.
Published: (2025)
ASTPrompter: Preference-Aligned Automated Language Model Red-Teaming to Generate Low-Perplexity Unsafe Prompts
by: Hardy, Amelia F., et al.
Published: (2024)
by: Hardy, Amelia F., et al.
Published: (2024)
Minimum Viable Ethics: From Institutionalizing Industry AI Governance to Product Impact
by: Ahlawat, Archana, et al.
Published: (2024)
by: Ahlawat, Archana, et al.
Published: (2024)
It's only fair when I think it's fair: How Gender Bias Alignment Undermines Distributive Fairness in Human-AI Collaboration
by: Zipperling, Domenique, et al.
Published: (2025)
by: Zipperling, Domenique, et al.
Published: (2025)
Can LLMs Synthesize Court-Ready Statistical Evidence? Evaluating AI-Assisted Sentencing Bias Analysis for California Racial Justice Act Claims
by: Komarla, Aparna
Published: (2026)
by: Komarla, Aparna
Published: (2026)
Take It, Leave It, or Fix It: Measuring Productivity and Trust in Human-AI Collaboration
by: Qian, Crystal, et al.
Published: (2024)
by: Qian, Crystal, et al.
Published: (2024)
The Laziness of the Crowd: Effort Aversion Among Raters Risks Undermining the Efficacy of X's Community Notes Program
by: Wack, Morgan, et al.
Published: (2026)
by: Wack, Morgan, et al.
Published: (2026)
Using Case Studies to Teach Responsible AI to Industry Practitioners
by: Stoyanovich, Julia, et al.
Published: (2024)
by: Stoyanovich, Julia, et al.
Published: (2024)
The Imbalanced User-AI Relationships as an Ethical Failure of Front-End Design in Healthcare AI
by: Mwadime, Maureen Mghambi
Published: (2026)
by: Mwadime, Maureen Mghambi
Published: (2026)
Kami of the Commons: Towards Designing Agentic AI to Steward the Commons
by: Hu, Botao Amber
Published: (2026)
by: Hu, Botao Amber
Published: (2026)
Bangladesh AI Readiness: Perspectives from the Academia, Industry, and Government
by: Sultana, Sharifa, et al.
Published: (2026)
by: Sultana, Sharifa, et al.
Published: (2026)
Examining Risks in the AI Companion Application Ecosystem
by: Brigham, Natalie Grace, et al.
Published: (2026)
by: Brigham, Natalie Grace, et al.
Published: (2026)
Exploring Multidimensional Checkworthiness: Designing AI-assisted Claim Prioritization for Human Fact-checkers
by: Liu, Houjiang, et al.
Published: (2024)
by: Liu, Houjiang, et al.
Published: (2024)
Rethinking AI-Mediated Minority Support in Power-Imbalanced Group Decision-Making: From Anonymity To Authenticity
by: Lee, Soohwan, et al.
Published: (2026)
by: Lee, Soohwan, et al.
Published: (2026)
Explainable Biomedical Claim Verification with Large Language Models
by: Liang, Siting, et al.
Published: (2025)
by: Liang, Siting, et al.
Published: (2025)
Agentic Visualization: Extracting Agent-based Design Patterns from Visualization Systems
by: Dhanoa, Vaishali, et al.
Published: (2025)
by: Dhanoa, Vaishali, et al.
Published: (2025)
Queering AI: Undoing the Self in the Algorithmic Borderlands
by: Turtle, Grace Leonora, et al.
Published: (2024)
by: Turtle, Grace Leonora, et al.
Published: (2024)
EyeAgent: An Agentic AI System for Multimodal Clinical Decision Support in Ophthalmology
by: Shi, Danli, et al.
Published: (2025)
by: Shi, Danli, et al.
Published: (2025)
Reassurance Robots: OCD in the Age of Generative AI
by: Barkhuff, Grace
Published: (2026)
by: Barkhuff, Grace
Published: (2026)
Agentic Enterprise: AI-Centric User to User-Centric AI
by: Narechania, Arpit, et al.
Published: (2025)
by: Narechania, Arpit, et al.
Published: (2025)
Why Johnny Can't Use Agents: Industry Aspirations vs. User Realities with AI Agents
by: Shome, Pradyumna, et al.
Published: (2025)
by: Shome, Pradyumna, et al.
Published: (2025)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
by: Vaccaro, Michelle, et al.
Published: (2026)
by: Vaccaro, Michelle, et al.
Published: (2026)
Imagining Design Workflows in Agentic AI Futures
by: Wadinambiarachchi, Samangi, et al.
Published: (2025)
by: Wadinambiarachchi, Samangi, et al.
Published: (2025)
A Practical Evaluation of Commercial Industrial Augmented Reality Systems in an Industry 4.0 Shipyard
by: Blanco-Novoa, Oscar, et al.
Published: (2024)
by: Blanco-Novoa, Oscar, et al.
Published: (2024)
Agentic AI: The Era of Semantic Decoding
by: Peyrard, Maxime, et al.
Published: (2024)
by: Peyrard, Maxime, et al.
Published: (2024)
Agentic AI as Undercover Teammates: Argumentative Knowledge Construction in Hybrid Human-AI Collaborative Learning
by: Yan, Lixiang, et al.
Published: (2025)
by: Yan, Lixiang, et al.
Published: (2025)
Agentic AI and Human-in-the-Loop Interventions: Field Experimental Evidence from Alibaba's Customer Service Operations
by: Wang, Yiwei, et al.
Published: (2026)
by: Wang, Yiwei, et al.
Published: (2026)
When Should Users Check? Modeling Confirmation Frequency inMulti-Step Agentic AI Tasks
by: Zhou, Jieyu, et al.
Published: (2025)
by: Zhou, Jieyu, et al.
Published: (2025)
Agentic Workflows for Conversational Human-AI Interaction Design
by: Caetano, Arthur, et al.
Published: (2025)
by: Caetano, Arthur, et al.
Published: (2025)
(AI peers) are people learning from the same standpoint: Perception of AI characters in a Collaborative Science Investigation
by: Ko, Eunhye Grace, et al.
Published: (2025)
by: Ko, Eunhye Grace, et al.
Published: (2025)
Similar Items
-
The Doctor Will (Still) See You Now: On the Structural Limits of Agentic AI in Healthcare
by: Dias, Gabriela Aránguiz, et al.
Published: (2026) -
An Adaptive Responsible AI Governance Framework for Decentralized Organizations
by: Meimandi, Kiana Jafari, et al.
Published: (2025) -
Responsible AI in the Global Context: Maturity Model and Survey
by: Reuel, Anka, et al.
Published: (2024) -
Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing
by: Jafari, Kiana, et al.
Published: (2026) -
More than Marketing? On the Information Value of AI Benchmarks for Practitioners
by: Hardy, Amelia, et al.
Published: (2024)