AgentRx: A Benchmark Study of LLM Agents for Multimodal Clinical Prediction Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Jorf, Baraa Al, Shamout, Farah E. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MedPatch: Confidence-Guided Multi-Stage Fusion for Multimodal Clinical Data
by: Jorf, Baraa Al, et al.
Published: (2025)
by: Jorf, Baraa Al, et al.
Published: (2025)
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
by: Barke, Shraddha, et al.
Published: (2026)
by: Barke, Shraddha, et al.
Published: (2026)
MIND: Modality-Informed Knowledge Distillation Framework for Multimodal Clinical Prediction Tasks
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025)
Uncertainty Quantification for Machine Learning in Healthcare: A Survey
by: López, L. Julián Lechuga, et al.
Published: (2025)
by: López, L. Julián Lechuga, et al.
Published: (2025)
MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
by: Daoud, Mouath Abu, et al.
Published: (2025)
by: Daoud, Mouath Abu, et al.
Published: (2025)
EHR-RAGp: Retrieval-Augmented Prototype-Guided Foundation Model for Electronic Health Records
by: Shurrab, Saeed, et al.
Published: (2026)
by: Shurrab, Saeed, et al.
Published: (2026)
RxEval: A Prescription-Level Benchmark for Evaluating LLM Medication Recommendation
by: Chen, Shuhao, et al.
Published: (2026)
by: Chen, Shuhao, et al.
Published: (2026)
NEWSAGENT: Benchmarking Multimodal Agents as Journalists with Real-World Newswriting Tasks
by: Chien, Yen-Che, et al.
Published: (2025)
by: Chien, Yen-Che, et al.
Published: (2025)
SafeRx-Agent: A Knowledge-Grounded Multi-Agent Framework for Safe and Explainable Medication Recommendation
by: Wang, Xinyu, et al.
Published: (2026)
by: Wang, Xinyu, et al.
Published: (2026)
Multi-modal Masked Siamese Network Improves Chest X-Ray Representation Learning
by: Shurrab, Saeed, et al.
Published: (2024)
by: Shurrab, Saeed, et al.
Published: (2024)
SafeAgentBench: A Benchmark for Safe Task Planning of Embodied LLM Agents
by: Yin, Sheng, et al.
Published: (2024)
by: Yin, Sheng, et al.
Published: (2024)
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
by: Yang, Xiao, et al.
Published: (2025)
by: Yang, Xiao, et al.
Published: (2025)
A Framework for Benchmarking and Aligning Task-Planning Safety in LLM-Based Embodied Agents
by: Huang, Yuting, et al.
Published: (2025)
by: Huang, Yuting, et al.
Published: (2025)
The Role of Functional Muscle Networks in Improving Hand Gesture Perception for Human-Machine Interfaces
by: Armanini, Costanza, et al.
Published: (2024)
by: Armanini, Costanza, et al.
Published: (2024)
Human-Guided Agentic AI for Multimodal Clinical Prediction: Lessons from the AgentDS Healthcare Benchmark
by: Pulavarthy, Lalitha Pranathi, et al.
Published: (2026)
by: Pulavarthy, Lalitha Pranathi, et al.
Published: (2026)
HWE-Bench: Benchmarking LLM Agents on Real-World Hardware Bug Repair Tasks
by: Cui, Fan, et al.
Published: (2026)
by: Cui, Fan, et al.
Published: (2026)
NeuroAgent: LLM Agents for Multimodal Neuroimaging Analysis and Research
by: Zhong, Lujia, et al.
Published: (2026)
by: Zhong, Lujia, et al.
Published: (2026)
CRAB: Cross-environment Agent Benchmark for Multimodal Language Model Agents
by: Xu, Tianqi, et al.
Published: (2024)
by: Xu, Tianqi, et al.
Published: (2024)
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
by: Xie, Tianbao, et al.
Published: (2024)
by: Xie, Tianbao, et al.
Published: (2024)
COMMA: A Communicative Multimodal Multi-Agent Benchmark
by: Ossowski, Timothy, et al.
Published: (2024)
by: Ossowski, Timothy, et al.
Published: (2024)
AUTOCT: Automating Interpretable Clinical Trial Prediction with LLM Agents
by: Liu, Fengze, et al.
Published: (2025)
by: Liu, Fengze, et al.
Published: (2025)
MCPAgentBench: A Real-world Task Benchmark for Evaluating LLM Agent MCP Tool Use
by: Liu, Wenrui, et al.
Published: (2025)
by: Liu, Wenrui, et al.
Published: (2025)
AgentLAB: Benchmarking LLM Agents against Long-Horizon Attacks
by: Jiang, Tanqiu, et al.
Published: (2026)
by: Jiang, Tanqiu, et al.
Published: (2026)
ReFuGe: Feature Generation for Prediction Tasks on Relational Databases with LLM Agents
by: Kim, Kyungho, et al.
Published: (2026)
by: Kim, Kyungho, et al.
Published: (2026)
Benchmarking LLM Agents for Wealth-Management Workflows
by: Milsom, Rory
Published: (2025)
by: Milsom, Rory
Published: (2025)
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
by: Zeng, Zhiyuan, et al.
Published: (2025)
by: Zeng, Zhiyuan, et al.
Published: (2025)
RewardHackingAgents: Benchmarking Evaluation Integrity for LLM ML-Engineering Agents
by: Atinafu, Yonas, et al.
Published: (2026)
by: Atinafu, Yonas, et al.
Published: (2026)
OSUniverse: Benchmark for Multimodal GUI-navigation AI Agents
by: Davydova, Mariya, et al.
Published: (2025)
by: Davydova, Mariya, et al.
Published: (2025)
VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications
by: He, Wei, et al.
Published: (2025)
by: He, Wei, et al.
Published: (2025)
Evaluation and Benchmarking of LLM Agents: A Survey
by: Mohammadi, Mahmoud, et al.
Published: (2025)
by: Mohammadi, Mahmoud, et al.
Published: (2025)
TerminalWorld: Benchmarking Agents on Real-World Terminal Tasks
by: Chu, Zhaoyang, et al.
Published: (2026)
by: Chu, Zhaoyang, et al.
Published: (2026)
AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Experience Transfer for Multimodal LLM Agents in Minecraft Game
by: Li, Chenghao, et al.
Published: (2026)
by: Li, Chenghao, et al.
Published: (2026)
ClinNoteAgents: An LLM Multi-Agent System for Predicting and Interpreting Heart Failure 30-Day Readmission from Clinical Notes
by: Zhou, Rongjia, et al.
Published: (2025)
by: Zhou, Rongjia, et al.
Published: (2025)
RiOSWorld: Benchmarking the Risk of Multimodal Computer-Use Agents
by: Yang, Jingyi, et al.
Published: (2025)
by: Yang, Jingyi, et al.
Published: (2025)
INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
by: Li, Haohang, et al.
Published: (2024)
by: Li, Haohang, et al.
Published: (2024)
Benchmarking LLM Summaries of Multimodal Clinical Time Series for Remote Monitoring
by: Shukla, Aditya, et al.
Published: (2026)
by: Shukla, Aditya, et al.
Published: (2026)
Ambig-DS: A Benchmark for Task-Framing Ambiguity in Data-Science Agents
by: Stoisser, Josefa Lia, et al.
Published: (2026)
by: Stoisser, Josefa Lia, et al.
Published: (2026)
Can LLM Agents Solve Collaborative Tasks? A Study on Urgency-Aware Planning and Coordination
by: Silva, João Vitor de Carvalho, et al.
Published: (2025)
by: Silva, João Vitor de Carvalho, et al.
Published: (2025)
Plan Verification for LLM-Based Embodied Task Completion Agents
by: Hariharan, Ananth, et al.
Published: (2025)
by: Hariharan, Ananth, et al.
Published: (2025)
Similar Items
-
MedPatch: Confidence-Guided Multi-Stage Fusion for Multimodal Clinical Data
by: Jorf, Baraa Al, et al.
Published: (2025) -
AgentRx: Diagnosing AI Agent Failures from Execution Trajectories
by: Barke, Shraddha, et al.
Published: (2026) -
MIND: Modality-Informed Knowledge Distillation Framework for Multimodal Clinical Prediction Tasks
by: Guerra-Manzanares, Alejandro, et al.
Published: (2025) -
Uncertainty Quantification for Machine Learning in Healthcare: A Survey
by: López, L. Julián Lechuga, et al.
Published: (2025) -
MedArabiQ: Benchmarking Large Language Models on Arabic Medical Tasks
by: Daoud, Mouath Abu, et al.
Published: (2025)