Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
Fuente:
arXiv
Saved in:
| Main Authors: | Kozlova, Anna, Salavei, Stanislau, Satalkin, Pavel, Plotnitskaya, Hanna, Parfenyuk, Sergey |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026)
by: Kong, Fanheng, et al.
Published: (2026)
Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
by: Yu, Ye, et al.
Published: (2026)
by: Yu, Ye, et al.
Published: (2026)
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
by: Sun, Yu, et al.
Published: (2025)
by: Sun, Yu, et al.
Published: (2025)
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
by: Zhu, Yinghao, et al.
Published: (2025)
by: Zhu, Yinghao, et al.
Published: (2025)
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
by: Jiang, Yixing, et al.
Published: (2025)
by: Jiang, Yixing, et al.
Published: (2025)
What if Pinocchio Were a Reinforcement Learning Agent: A Normative End-to-End Pipeline
by: Alcaraz, Benoît
Published: (2026)
by: Alcaraz, Benoît
Published: (2026)
MedPriv-Bench: Benchmarking the Privacy-Utility Trade-off of Large Language Models in Medical Open-End Question Answering
by: Guan, Shaowei, et al.
Published: (2026)
by: Guan, Shaowei, et al.
Published: (2026)
MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
by: Fang, Haoyang, et al.
Published: (2025)
by: Fang, Haoyang, et al.
Published: (2025)
Dialogue Diplomats: An End-to-End Multi-Agent Reinforcement Learning System for Automated Conflict Resolution and Consensus Building
by: Bolleddu, Deepak
Published: (2025)
by: Bolleddu, Deepak
Published: (2025)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
by: Zhu, Kunlun, et al.
Published: (2025)
by: Zhu, Kunlun, et al.
Published: (2025)
WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
by: Styles, Olly, et al.
Published: (2024)
by: Styles, Olly, et al.
Published: (2024)
DDO: Dual-Decision Optimization for LLM-Based Medical Consultation via Multi-Agent Collaboration
by: Jia, Zhihao, et al.
Published: (2025)
by: Jia, Zhihao, et al.
Published: (2025)
MRGAgents: A Multi-Agent Framework for Improved Medical Report Generation with Med-LVLMs
by: Wang, Pengyu, et al.
Published: (2025)
by: Wang, Pengyu, et al.
Published: (2025)
MedSentry: Understanding and Mitigating Safety Risks in Medical LLM Multi-Agent Systems
by: Chen, Kai, et al.
Published: (2025)
by: Chen, Kai, et al.
Published: (2025)
CORE-Bench: Fostering the Credibility of Published Research Through a Computational Reproducibility Agent Benchmark
by: Siegel, Zachary S., et al.
Published: (2024)
by: Siegel, Zachary S., et al.
Published: (2024)
CACA Agent: Capability Collaboration based AI Agent
by: Xu, Peng, et al.
Published: (2024)
by: Xu, Peng, et al.
Published: (2024)
ConfAgents: A Conformal-Guided Multi-Agent Framework for Cost-Efficient Medical Diagnosis
by: Zhao, Huiya, et al.
Published: (2025)
by: Zhao, Huiya, et al.
Published: (2025)
IryoNLP at MEDIQA-CORR 2024: Tackling the Medical Error Detection & Correction Task On the Shoulders of Medical Agents
by: Corbeil, Jean-Philippe
Published: (2024)
by: Corbeil, Jean-Philippe
Published: (2024)
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
by: Bogavelli, Tara, et al.
Published: (2025)
by: Bogavelli, Tara, et al.
Published: (2025)
MedRAX: Medical Reasoning Agent for Chest X-ray
by: Fallahpour, Adibvafa, et al.
Published: (2025)
by: Fallahpour, Adibvafa, et al.
Published: (2025)
UCAgent: An End-to-End Agent for Block-Level Functional Verification
by: Wang, Junyue, et al.
Published: (2026)
by: Wang, Junyue, et al.
Published: (2026)
MedDCR: Learning to Design Agentic Workflows for Medical Coding
by: Zheng, Jiyang, et al.
Published: (2025)
by: Zheng, Jiyang, et al.
Published: (2025)
Every 28 Days the AI Dreams of Soft Skin and Burning Stars: Scaffolding AI Agents with Hormones and Emotions
by: Levinson, Leigh, et al.
Published: (2025)
by: Levinson, Leigh, et al.
Published: (2025)
LiteWebAgent: The Open-Source Suite for VLM-Based Web-Agent Applications
by: Zhang, Danqing, et al.
Published: (2025)
by: Zhang, Danqing, et al.
Published: (2025)
Prompt Injection Detection and Mitigation via AI Multi-Agent NLP Frameworks
by: Gosmar, Diego, et al.
Published: (2025)
by: Gosmar, Diego, et al.
Published: (2025)
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
by: Sun, Haochen, et al.
Published: (2025)
by: Sun, Haochen, et al.
Published: (2025)
Adaptive Monitoring and Real-World Evaluation of Agentic AI Systems
by: Shukla, Manish
Published: (2025)
by: Shukla, Manish
Published: (2025)
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
by: Zou, Chelsea, et al.
Published: (2026)
by: Zou, Chelsea, et al.
Published: (2026)
LC-Opt: Benchmarking Reinforcement Learning and Agentic AI for End-to-End Liquid Cooling Optimization in Data Centers
by: Naug, Avisek, et al.
Published: (2025)
by: Naug, Avisek, et al.
Published: (2025)
Agents on the Bench: Large Language Model Based Multi Agent Framework for Trustworthy Digital Justice
by: Jiang, Cong, et al.
Published: (2024)
by: Jiang, Cong, et al.
Published: (2024)
AMANDA: Agentic Medical Knowledge Augmentation for Data-Efficient Medical Visual Question Answering
by: Wang, Ziqing, et al.
Published: (2025)
by: Wang, Ziqing, et al.
Published: (2025)
AssetOpsBench: Benchmarking AI Agents for Task Automation in Industrial Asset Operations and Maintenance
by: Patel, Dhaval, et al.
Published: (2025)
by: Patel, Dhaval, et al.
Published: (2025)
Hallucination Mitigation using Agentic AI Natural Language-Based Frameworks
by: Gosmar, Diego, et al.
Published: (2025)
by: Gosmar, Diego, et al.
Published: (2025)
Manalyzer: End-to-end Automated Meta-analysis with Multi-agent System
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
RPA-Check: A Multi-Stage Automated Framework for Evaluating Dynamic LLM-based Role-Playing Agents
by: Rosati, Riccardo, et al.
Published: (2026)
by: Rosati, Riccardo, et al.
Published: (2026)
AlphaDou: High-Performance End-to-End Doudizhu AI Integrating Bidding
by: Lei, Chang, et al.
Published: (2024)
by: Lei, Chang, et al.
Published: (2024)
Silo-Bench: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
by: Zhang, Yuzhe, et al.
Published: (2026)
by: Zhang, Yuzhe, et al.
Published: (2026)
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
by: Fang, Jinyuan, et al.
Published: (2025)
by: Fang, Jinyuan, et al.
Published: (2025)
MARBLE: A Multi-Agent Rule-Based LLM Reasoning Engine for Accident Severity Prediction
by: Qasim, Kaleem Ullah, et al.
Published: (2025)
by: Qasim, Kaleem Ullah, et al.
Published: (2025)
Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems
by: Zhou, Ruiwen, et al.
Published: (2026)
by: Zhou, Ruiwen, et al.
Published: (2026)
Similar Items
-
WebTestBench: Evaluating Computer-Use Agents towards End-to-End Automated Web Testing
by: Kong, Fanheng, et al.
Published: (2026) -
Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems
by: Yu, Ye, et al.
Published: (2026) -
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
by: Sun, Yu, et al.
Published: (2025) -
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
by: Zhu, Yinghao, et al.
Published: (2025) -
MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
by: Jiang, Yixing, et al.
Published: (2025)