Measuring What Matters: Benchmarking Generative, Multimodal, and Agentic AI in Healthcare
Fuente:
arXiv
Saved in:
| Main Authors: | Desikan, Prasanna, Rajgarhia, Harshit, Dalmia, Shivali, Mantravadi, Ananya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction
by: Dalmia, Shivali, et al.
Published: (2026)
by: Dalmia, Shivali, et al.
Published: (2026)
ART: Action-based Reasoning Task Benchmarking for Medical AI Agents
by: Mantravadi, Ananya, et al.
Published: (2026)
by: Mantravadi, Ananya, et al.
Published: (2026)
Human + AI for Accelerating Ad Localization Evaluation
by: Rajgarhia, Harshit, et al.
Published: (2025)
by: Rajgarhia, Harshit, et al.
Published: (2025)
LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents
by: Mantravadi, Ananya, et al.
Published: (2025)
by: Mantravadi, Ananya, et al.
Published: (2025)
Scalable multilingual PII annotation for responsible AI in LLMs
by: Meena, Bharti, et al.
Published: (2025)
by: Meena, Bharti, et al.
Published: (2025)
MedMosaic: A Challenging Large Scale Benchmark of Diverse Medical Audio
by: Rajgarhia, Harshit, et al.
Published: (2026)
by: Rajgarhia, Harshit, et al.
Published: (2026)
Automated Auditing of Hospital Discharge Summaries for Care Transitions
by: Dasula, Akshat, et al.
Published: (2026)
by: Dasula, Akshat, et al.
Published: (2026)
Enhancing Guardrails for Safe and Secure Healthcare AI
by: Gangavarapu, Ananya
Published: (2024)
by: Gangavarapu, Ananya
Published: (2024)
GAZE:Governance-Aware pre-annotation for Zero-shot World Model Environments
by: Krishna, Leela, et al.
Published: (2025)
by: Krishna, Leela, et al.
Published: (2025)
IMAS: A Comprehensive Agentic Approach to Rural Healthcare Delivery
by: Gangavarapu, Agasthya, et al.
Published: (2024)
by: Gangavarapu, Agasthya, et al.
Published: (2024)
Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark
by: Pulavarthy, Lalitha Pranathi, et al.
Published: (2026)
by: Pulavarthy, Lalitha Pranathi, et al.
Published: (2026)
Measuring What Matters: The AI Pluralism Index
by: Mushkani, Rashid
Published: (2025)
by: Mushkani, Rashid
Published: (2025)
AI Safeguards, Generative AI and the Pandora Box: AI Safety Measures to Protect Businesses and Personal Reputation
by: Kumar, Prasanna
Published: (2026)
by: Kumar, Prasanna
Published: (2026)
Agentic AI Scientists Are Not Built For Autonomous Scientific Discovery
by: Bisht, Harshit, et al.
Published: (2026)
by: Bisht, Harshit, et al.
Published: (2026)
Detecting Silent Failures in Multi-Agentic AI Trajectories
by: Pathak, Divya, et al.
Published: (2025)
by: Pathak, Divya, et al.
Published: (2025)
Agentic Temporal Graph of Reasoning with Multimodal Language Models: A Potential AI Aid to Healthcare
by: Mitra, Susanta
Published: (2025)
by: Mitra, Susanta
Published: (2025)
Agentic AI Governance and Lifecycle Management in Healthcare
by: Prakash, Chandra, et al.
Published: (2026)
by: Prakash, Chandra, et al.
Published: (2026)
An Evaluation Study of Hybrid Methods for Multilingual PII Detection
by: Rajgarhia, Harshit, et al.
Published: (2025)
by: Rajgarhia, Harshit, et al.
Published: (2025)
Agentic AI for Self-Driving Laboratories in Soft Matter: Taxonomy, Benchmarks,and Open Challenges
by: Chen, Xuanzhou, et al.
Published: (2026)
by: Chen, Xuanzhou, et al.
Published: (2026)
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?
by: Wei, Qianshan, et al.
Published: (2026)
by: Wei, Qianshan, et al.
Published: (2026)
Measuring What Matters Beyond Text: Evaluating Multimodal Summaries by Quality, Alignment, and Diversity
by: Ali, Abid, et al.
Published: (2026)
by: Ali, Abid, et al.
Published: (2026)
Constrained Process Maps for Multi-Agent Generative AI Workflows
by: Joshi, Ananya, et al.
Published: (2026)
by: Joshi, Ananya, et al.
Published: (2026)
Benchmarking Edge AI Platforms for High-Performance ML Inference
by: Jayanth, Rakshith, et al.
Published: (2024)
by: Jayanth, Rakshith, et al.
Published: (2024)
Notation Matters: A Benchmark Study of Token-Optimized Formats in Agentic AI Systems
by: Kutschka, Lorenz, et al.
Published: (2026)
by: Kutschka, Lorenz, et al.
Published: (2026)
Grounded by Experience: Generative Healthcare Prediction Augmented with Hierarchical Agentic Retrieval
by: Zhao, Chuang, et al.
Published: (2025)
by: Zhao, Chuang, et al.
Published: (2025)
SWEnergy: An Empirical Study on Energy Efficiency in Agentic Issue Resolution Frameworks with SLMs
by: Tripathy, Arihant, et al.
Published: (2025)
by: Tripathy, Arihant, et al.
Published: (2025)
Agentic AI Needs a Systems Theory
by: Miehling, Erik, et al.
Published: (2025)
by: Miehling, Erik, et al.
Published: (2025)
AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction
by: Mankodiya, Harsh, et al.
Published: (2026)
by: Mankodiya, Harsh, et al.
Published: (2026)
AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
by: Ferrag, Mohamed Amine, et al.
Published: (2026)
Towards a HIPAA Compliant Agentic AI System in Healthcare
by: Neupane, Subash, et al.
Published: (2025)
by: Neupane, Subash, et al.
Published: (2025)
Generative AI for Healthcare: Fundamentals, Challenges, and Perspectives
by: Chen, Gang, et al.
Published: (2025)
by: Chen, Gang, et al.
Published: (2025)
PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
Generative to Agentic AI: Survey, Conceptualization, and Challenges
by: Schneider, Johannes
Published: (2025)
by: Schneider, Johannes
Published: (2025)
UAVBench: An Open Benchmark Dataset for Autonomous and Agentic AI UAV Systems via LLM-Generated Flight Scenarios
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
by: Ferrag, Mohamed Amine, et al.
Published: (2025)
What Matters in Data Curation for Multimodal Reasoning? Insights from the DCVLR Challenge
by: Shin, Yosub, et al.
Published: (2026)
by: Shin, Yosub, et al.
Published: (2026)
Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
by: Guan, Zihan, et al.
Published: (2026)
by: Guan, Zihan, et al.
Published: (2026)
Measuring What Matters: Connecting AI Ethics Evaluations to System Attributes, Hazards, and Harms
by: Rismani, Shalaleh, et al.
Published: (2025)
by: Rismani, Shalaleh, et al.
Published: (2025)
Med-MMFL: A Multimodal Federated Learning Benchmark in Healthcare
by: Chhetri, Aavash, et al.
Published: (2026)
by: Chhetri, Aavash, et al.
Published: (2026)
OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents
by: Davydova, Mariya, et al.
Published: (2025)
by: Davydova, Mariya, et al.
Published: (2025)
StarBench: A Turn-Based RPG Benchmark for Agentic Multimodal Decision-Making and Information Seeking
by: Zhang, Haoran, et al.
Published: (2025)
by: Zhang, Haoran, et al.
Published: (2025)
Similar Items
-
Systematic Evaluation of Large Language Models for Post-Discharge Clinical Action Extraction
by: Dalmia, Shivali, et al.
Published: (2026) -
ART: Action-based Reasoning Task Benchmarking for Medical AI Agents
by: Mantravadi, Ananya, et al.
Published: (2026) -
Human + AI for Accelerating Ad Localization Evaluation
by: Rajgarhia, Harshit, et al.
Published: (2025) -
LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents
by: Mantravadi, Ananya, et al.
Published: (2025) -
Scalable multilingual PII annotation for responsible AI in LLMs
by: Meena, Bharti, et al.
Published: (2025)