HealthAdminBench: Evaluating Computer-Use Agents on Healthcare Administration Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Bedi, Suhana, Welch, Ryan, Steinberg, Ethan, Wornow, Michael, Kim, Taeil Matthew, Ahmed, Haroun, Sterling, Peter, Purohit, Bravim, Akram, Qurat, Acosta, Angelic, Nubla, Esther, Sharma, Pritika, Pfeffer, Michael A., Koyejo, Sanmi, Shah, Nigam H. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHRs
by: Wornow, Michael, et al.
Published: (2024)
by: Wornow, Michael, et al.
Published: (2024)
The Optimization Paradox in Clinical AI Multi-Agent Systems
by: Bedi, Suhana, et al.
Published: (2025)
by: Bedi, Suhana, et al.
Published: (2025)
meds_reader: A fast and efficient EHR processing library
by: Steinberg, Ethan, et al.
Published: (2024)
by: Steinberg, Ethan, et al.
Published: (2024)
Quantifying and Mitigating Premature Closure in Frontier LLMs
by: Handler, Rebecca, et al.
Published: (2026)
by: Handler, Rebecca, et al.
Published: (2026)
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
by: Obbad, Elyas, et al.
Published: (2024)
by: Obbad, Elyas, et al.
Published: (2024)
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Salesforce Salesforce Administration Essentials for New Admins PDF
by: Certification Exam
Published: (2026)
by: Certification Exam
Published: (2026)
Salesforce Salesforce Administration Essentials for New Admins PDF
by: Certification Exam
Published: (2026)
by: Certification Exam
Published: (2026)
Distilling Large Language Models for Efficient Clinical Information Extraction
by: Vedula, Karthik S., et al.
Published: (2024)
by: Vedula, Karthik S., et al.
Published: (2024)
Structured Prompts Improve Evaluation of Language Models
by: Aali, Asad, et al.
Published: (2025)
by: Aali, Asad, et al.
Published: (2025)
Causally Inspired Regularization Enables Domain General Representations
by: Salaudeen, Olawale, et al.
Published: (2024)
by: Salaudeen, Olawale, et al.
Published: (2024)
Let's Measure Information Step-by-Step: AI-Based Evaluation Beyond Vibes
by: Robertson, Zachary, et al.
Published: (2025)
by: Robertson, Zachary, et al.
Published: (2025)
CURE: Cultural Understanding and Reasoning Evaluation - A Framework for "Thick" Culture Alignment Evaluation in LLMs
by: Vo, Truong, et al.
Published: (2025)
by: Vo, Truong, et al.
Published: (2025)
A Framework for Objective-Driven Dynamical Stochastic Fields
by: Zhang, Yibo Jacky, et al.
Published: (2025)
by: Zhang, Yibo Jacky, et al.
Published: (2025)
In-Context Learning of Energy Functions
by: Schaeffer, Rylan, et al.
Published: (2024)
by: Schaeffer, Rylan, et al.
Published: (2024)
Discovering Implicit Large Language Model Alignment Objectives
by: Chen, Edward, et al.
Published: (2026)
by: Chen, Edward, et al.
Published: (2026)
High-Dimensional Markov-switching Ordinary Differential Processes
by: Tsai, Katherine, et al.
Published: (2024)
by: Tsai, Katherine, et al.
Published: (2024)
Distributional Machine Unlearning via Selective Data Removal
by: Allouah, Youssef, et al.
Published: (2025)
by: Allouah, Youssef, et al.
Published: (2025)
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases
by: Iyer, Laya, et al.
Published: (2026)
by: Iyer, Laya, et al.
Published: (2026)
HiFA: High-fidelity Text-to-3D Generation with Advanced Diffusion Guidance
by: Zhu, Junzhe, et al.
Published: (2023)
by: Zhu, Junzhe, et al.
Published: (2023)
Selection of filamentous fungi that are resistant to the herbicides atrazine, glyphosate and pendimethalin
by: Nara Priscila Barbosa Bravim
Published: (2021)
by: Nara Priscila Barbosa Bravim
Published: (2021)
Zero-Shot Clinical Trial Patient Matching with LLMs
by: Wornow, Michael, et al.
Published: (2024)
by: Wornow, Michael, et al.
Published: (2024)
TIMER: Temporal Instruction Modeling and Evaluation for Longitudinal Clinical Records
by: Cui, Hejie, et al.
Published: (2025)
by: Cui, Hejie, et al.
Published: (2025)
Reasoning Models Don't Just Think Longer, They Move Differently
by: Gjølbye, Anders, et al.
Published: (2026)
by: Gjølbye, Anders, et al.
Published: (2026)
Is Backpropagation Optimal? When Synthetic Gradients Improve Sample Efficiency
by: Zhang, Yibo Jacky, et al.
Published: (2026)
by: Zhang, Yibo Jacky, et al.
Published: (2026)
The Inadequacy of Offline LLM Evaluations: A Need to Account for Personalization in Model Behavior
by: Wang, Angelina, et al.
Published: (2025)
by: Wang, Angelina, et al.
Published: (2025)
Principled Federated Domain Adaptation: Gradient Projection and Auto-Weighting
by: Jiang, Enyi, et al.
Published: (2023)
by: Jiang, Enyi, et al.
Published: (2023)
H-AdminSim: A Multi-Agent Simulator for Realistic Hospital Administrative Workflows with FHIR Integration
by: Lee, Jun-Min, et al.
Published: (2026)
by: Lee, Jun-Min, et al.
Published: (2026)
Beyond Buttons: Is the Admin Dashboard Dead?
by: Rosehill, Daniel, et al.
Published: (2026)
by: Rosehill, Daniel, et al.
Published: (2026)
Automating the Enterprise with Foundation Models
by: Wornow, Michael, et al.
Published: (2024)
by: Wornow, Michael, et al.
Published: (2024)
Evolution, Challenges, and Future Prospects of Banking and Financial Institutions
by: Varshney, Pritika
Published: (2025)
by: Varshney, Pritika
Published: (2025)
Why Do Safety Guardrails Degrade Across Languages?
by: Zhang, Max, et al.
Published: (2026)
by: Zhang, Max, et al.
Published: (2026)
Are Domain Generalization Benchmarks with Accuracy on the Line Misspecified?
by: Salaudeen, Olawale, et al.
Published: (2025)
by: Salaudeen, Olawale, et al.
Published: (2025)
Pretraining Scaling Laws for Generative Evaluations of Language Models
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
The Utility and Complexity of in- and out-of-Distribution Machine Unlearning
by: Allouah, Youssef, et al.
Published: (2024)
by: Allouah, Youssef, et al.
Published: (2024)
Logits are All We Need to Adapt Closed Models
by: Hiranandani, Gaurush, et al.
Published: (2025)
by: Hiranandani, Gaurush, et al.
Published: (2025)
Invariant Aggregator for Defending against Federated Backdoor Attacks
by: Wang, Xiaoyang, et al.
Published: (2022)
by: Wang, Xiaoyang, et al.
Published: (2022)
Stop Automating Peer Review Without Rigorous Evaluation
by: Baumann, Joachim, et al.
Published: (2026)
by: Baumann, Joachim, et al.
Published: (2026)
Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM Agents
by: Zhang, Qizheng, et al.
Published: (2025)
by: Zhang, Qizheng, et al.
Published: (2025)
Fairness through Difference Awareness: Measuring Desired Group Discrimination in LLMs
by: Wang, Angelina, et al.
Published: (2025)
by: Wang, Angelina, et al.
Published: (2025)
Similar Items
-
Context Clues: Evaluating Long Context Models for Clinical Prediction Tasks on EHRs
by: Wornow, Michael, et al.
Published: (2024) -
The Optimization Paradox in Clinical AI Multi-Agent Systems
by: Bedi, Suhana, et al.
Published: (2025) -
meds_reader: A fast and efficient EHR processing library
by: Steinberg, Ethan, et al.
Published: (2024) -
Quantifying and Mitigating Premature Closure in Frontier LLMs
by: Handler, Rebecca, et al.
Published: (2026) -
ZIP-FIT: Embedding-Free Data Selection via Compression-Based Alignment
by: Obbad, Elyas, et al.
Published: (2024)