MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jiang, Yixing, Black, Kameron C., Geng, Gloria, Park, Danny, Zou, James, Ng, Andrew Y., Chen, Jonathan H. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ALAS: Transactional and Dynamic Multi-Agent LLM Planning
von: Geng, Longling, et al.
Veröffentlicht: (2025)
von: Geng, Longling, et al.
Veröffentlicht: (2025)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
von: Styles, Olly, et al.
Veröffentlicht: (2024)
von: Styles, Olly, et al.
Veröffentlicht: (2024)
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
von: Zhong, Shanshan, et al.
Veröffentlicht: (2026)
von: Zhong, Shanshan, et al.
Veröffentlicht: (2026)
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
von: Kozlova, Anna, et al.
Veröffentlicht: (2026)
von: Kozlova, Anna, et al.
Veröffentlicht: (2026)
MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems
von: Chen, Kai, et al.
Veröffentlicht: (2025)
von: Chen, Kai, et al.
Veröffentlicht: (2025)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Yuzhe, et al.
Veröffentlicht: (2026)
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
von: Zhu, Yinghao, et al.
Veröffentlicht: (2025)
von: Zhu, Yinghao, et al.
Veröffentlicht: (2025)
Safe Multi-Agent Behavior Must Be Maintained, Not Merely Asserted: Constraint Drift in LLM-Based Multi-Agent Systems
von: Li, Tianxiao, et al.
Veröffentlicht: (2026)
von: Li, Tianxiao, et al.
Veröffentlicht: (2026)
MedRAX: Medical Reasoning Agent for Chest X-ray
von: Fallahpour, Adibvafa, et al.
Veröffentlicht: (2025)
von: Fallahpour, Adibvafa, et al.
Veröffentlicht: (2025)
PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
von: Schipper, Olivier, et al.
Veröffentlicht: (2025)
von: Schipper, Olivier, et al.
Veröffentlicht: (2025)
SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
von: Hamblin, Grant, et al.
Veröffentlicht: (2026)
von: Hamblin, Grant, et al.
Veröffentlicht: (2026)
ClinEnv: An Interactive Multi-Stage Long Horizon EHR Environment for Agents
von: Lu, Yuxing, et al.
Veröffentlicht: (2026)
von: Lu, Yuxing, et al.
Veröffentlicht: (2026)
AgentSearchBench: A Benchmark for AI Agent Search in the Wild
von: Wu, Bin, et al.
Veröffentlicht: (2026)
von: Wu, Bin, et al.
Veröffentlicht: (2026)
MRGAgents: A Multi-Agent Framework for Improved Medical Report Generation with Med-LVLMs
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
von: Wang, Pengyu, et al.
Veröffentlicht: (2025)
LingxiDiagBench: A Multi-Agent Framework for Benchmarking LLMs in Chinese Psychiatric Consultation and Diagnosis
von: Xu, Shihao, et al.
Veröffentlicht: (2026)
von: Xu, Shihao, et al.
Veröffentlicht: (2026)
MLC-Agent: Cognitive Model based on Memory-Learning Collaboration in LLM Empowered Agent Simulation Environment
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
von: Zhang, Ming, et al.
Veröffentlicht: (2025)
MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering
von: Guan, Shaowei, et al.
Veröffentlicht: (2026)
von: Guan, Shaowei, et al.
Veröffentlicht: (2026)
Agents on the Bench: Large Language Model Based Multi Agent Framework for Trustworthy Digital Justice
von: Jiang, Cong, et al.
Veröffentlicht: (2024)
von: Jiang, Cong, et al.
Veröffentlicht: (2024)
S-Agents: Self-organizing Agents in Open-ended Environments
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
von: Chen, Jiaqi, et al.
Veröffentlicht: (2024)
BenchMARL: Benchmarking Multi-Agent Reinforcement Learning
von: Bettini, Matteo, et al.
Veröffentlicht: (2023)
von: Bettini, Matteo, et al.
Veröffentlicht: (2023)
Self-Refining Topology Optimization via an LLM-Based Multi-Agent Framework
von: Park, Hyunjee, et al.
Veröffentlicht: (2026)
von: Park, Hyunjee, et al.
Veröffentlicht: (2026)
SWE-WebDevBench: Evaluating Coding Agent Application Platforms as Virtual Software Agencies
von: Saxena, Siddhant, et al.
Veröffentlicht: (2026)
von: Saxena, Siddhant, et al.
Veröffentlicht: (2026)
PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments
von: Liu, Ruoqi, et al.
Veröffentlicht: (2026)
von: Liu, Ruoqi, et al.
Veröffentlicht: (2026)
Generalist Virtual Agents: A Survey on Autonomous Agents Across Digital Platforms
von: Gao, Minghe, et al.
Veröffentlicht: (2024)
von: Gao, Minghe, et al.
Veröffentlicht: (2024)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
von: Zhang, Boxuan, et al.
Veröffentlicht: (2025)
HSCodeComp: A Realistic and Expert-level Benchmark for Deep Search Agents in Hierarchical Rule Application
von: Yang, Yiqian, et al.
Veröffentlicht: (2025)
von: Yang, Yiqian, et al.
Veröffentlicht: (2025)
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
von: Zou, Chelsea, et al.
Veröffentlicht: (2026)
von: Zou, Chelsea, et al.
Veröffentlicht: (2026)
Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents
von: Song, Kevin, et al.
Veröffentlicht: (2025)
von: Song, Kevin, et al.
Veröffentlicht: (2025)
FinDeepForecast: A Live Multi-Agent System for Benchmarking Deep Research Agents in Financial Forecasting
von: Li, Xiangyu, et al.
Veröffentlicht: (2026)
von: Li, Xiangyu, et al.
Veröffentlicht: (2026)
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
von: Sun, Yu, et al.
Veröffentlicht: (2025)
von: Sun, Yu, et al.
Veröffentlicht: (2025)
Scaling Lifelong Multi-Agent Path Finding to More Realistic Settings: Research Challenges and Opportunities
von: Jiang, He, et al.
Veröffentlicht: (2024)
von: Jiang, He, et al.
Veröffentlicht: (2024)
AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
von: Patel, Dhaval, et al.
Veröffentlicht: (2025)
von: Patel, Dhaval, et al.
Veröffentlicht: (2025)
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
von: Kavathekar, Ishan, et al.
Veröffentlicht: (2025)
von: Kavathekar, Ishan, et al.
Veröffentlicht: (2025)
Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems
von: Chen, Mengzhuo, et al.
Veröffentlicht: (2026)
von: Chen, Mengzhuo, et al.
Veröffentlicht: (2026)
CREW-WILDFIRE: Benchmarking Agentic Multi-Agent Collaborations at Scale
von: Hyun, Jonathan, et al.
Veröffentlicht: (2025)
von: Hyun, Jonathan, et al.
Veröffentlicht: (2025)
Multi-Agent Team Access Monitoring: Environments that Benefit from Target Information Sharing
von: Dudash, Andrew, et al.
Veröffentlicht: (2024)
von: Dudash, Andrew, et al.
Veröffentlicht: (2024)
FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models
von: Ye, Rui, et al.
Veröffentlicht: (2024)
von: Ye, Rui, et al.
Veröffentlicht: (2024)
ArgMed-Agents: Explainable Clinical Decision Reasoning with LLM Disscusion via Argumentation Schemes
von: Hong, Shengxin, et al.
Veröffentlicht: (2024)
von: Hong, Shengxin, et al.
Veröffentlicht: (2024)
iAgentBench: Benchmarking Sensemaking Capabilities of Information-Seeking Agents on High-Traffic Topics
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2026)
von: Dammu, Preetam Prabhu Srikar, et al.
Veröffentlicht: (2026)
Tapas Are Free! Training-Free Adaptation of Programmatic Agents via LLM-Guided Program Synthesis in Dynamic Environments
von: Hu, Jinwei, et al.
Veröffentlicht: (2025)
von: Hu, Jinwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ALAS: Transactional and Dynamic Multi-Agent LLM Planning
von: Geng, Longling, et al.
Veröffentlicht: (2025) -
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
von: Styles, Olly, et al.
Veröffentlicht: (2024) -
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
von: Zhong, Shanshan, et al.
Veröffentlicht: (2026) -
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
von: Kozlova, Anna, et al.
Veröffentlicht: (2026) -
MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems
von: Chen, Kai, et al.
Veröffentlicht: (2025)