HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bedi, Suhana, Welch, Ryan, Steinberg, Ethan, Wornow, Michael, Kim, Taeil Matthew, Ahmed, Haroun, Sterling, Peter, Purohit, Bravim, Akram, Qurat, Acosta, Angelic, Nubla, Esther, Sharma, Pritika, Pfeffer, Michael A., Koyejo, Sanmi, Shah, Nigam H.
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911584761151488
author Bedi, Suhana
Welch, Ryan
Steinberg, Ethan
Wornow, Michael
Kim, Taeil Matthew
Ahmed, Haroun
Sterling, Peter
Purohit, Bravim
Akram, Qurat
Acosta, Angelic
Nubla, Esther
Sharma, Pritika
Pfeffer, Michael A.
Koyejo, Sanmi
Shah, Nigam H.
author_facet Bedi, Suhana
Welch, Ryan
Steinberg, Ethan
Wornow, Michael
Kim, Taeil Matthew
Ahmed, Haroun
Sterling, Peter
Purohit, Bravim
Akram, Qurat
Acosta, Angelic
Nubla, Esther
Sharma, Pritika
Pfeffer, Michael A.
Koyejo, Sanmi
Shah, Nigam H.
contents Healthcare administration accounts for over $1 trillion in annual spending, making it a promising target for LLM-based computer-use agents (CUAs). While clinical applications of LLMs have received significant attention, no benchmark exists for evaluating CUAs on end-to-end administrative workflows. To address this gap, we introduce HealthAdminBench, a benchmark comprising four realistic GUI environments: an EHR, two payer portals, and a fax system, and 135 expert-defined tasks spanning three administrative task types: Prior Authorization, Appeals and Denials Management, and Durable Medical Equipment (DME) Order Processing. Each task is decomposed into fine-grained, verifiable subtasks, yielding 1,698 evaluation points. We evaluate seven agent configurations under multiple prompting and observation settings and find that, despite strong subtask performance, end-to-end reliability remains low: the best-performing agent (Claude Opus 4.6 CUA) achieves only 36.3 percent task success, while GPT-5.4 CUA attains the highest subtask success rate (82.8 percent). These results reveal a substantial gap between current agent capabilities and the demands of real-world administrative workflows. HealthAdminBench provides a rigorous foundation for evaluating progress toward safe and reliable automation of healthcare administrative workflows.
format Preprint
id arxiv_https___arxiv_org_abs_2604_09937
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
Bedi, Suhana
Welch, Ryan
Steinberg, Ethan
Wornow, Michael
Kim, Taeil Matthew
Ahmed, Haroun
Sterling, Peter
Purohit, Bravim
Akram, Qurat
Acosta, Angelic
Nubla, Esther
Sharma, Pritika
Pfeffer, Michael A.
Koyejo, Sanmi
Shah, Nigam H.
Artificial Intelligence
Healthcare administration accounts for over $1 trillion in annual spending, making it a promising target for LLM-based computer-use agents (CUAs). While clinical applications of LLMs have received significant attention, no benchmark exists for evaluating CUAs on end-to-end administrative workflows. To address this gap, we introduce HealthAdminBench, a benchmark comprising four realistic GUI environments: an EHR, two payer portals, and a fax system, and 135 expert-defined tasks spanning three administrative task types: Prior Authorization, Appeals and Denials Management, and Durable Medical Equipment (DME) Order Processing. Each task is decomposed into fine-grained, verifiable subtasks, yielding 1,698 evaluation points. We evaluate seven agent configurations under multiple prompting and observation settings and find that, despite strong subtask performance, end-to-end reliability remains low: the best-performing agent (Claude Opus 4.6 CUA) achieves only 36.3 percent task success, while GPT-5.4 CUA attains the highest subtask success rate (82.8 percent). These results reveal a substantial gap between current agent capabilities and the demands of real-world administrative workflows. HealthAdminBench provides a rigorous foundation for evaluating progress toward safe and reliable automation of healthcare administrative workflows.
title HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
topic Artificial Intelligence
url https://arxiv.org/abs/2604.09937