AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Gyutaek, Park, Sangjoon, Kim, Byung-Hoon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025)
by: Oh, Gyutaek, et al.
Published: (2025)
AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML
by: Trirat, Patara, et al.
Published: (2024)
by: Trirat, Patara, et al.
Published: (2024)
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
by: Kozlova, Anna, et al.
Published: (2026)
by: Kozlova, Anna, et al.
Published: (2026)
MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation
by: Wang, George, et al.
Published: (2025)
by: Wang, George, et al.
Published: (2025)
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
by: Zhu, Yinghao, et al.
Published: (2025)
by: Zhu, Yinghao, et al.
Published: (2025)
AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
by: Ou, Yixin, et al.
Published: (2025)
by: Ou, Yixin, et al.
Published: (2025)
PAACE: A Plan-Aware Automated Agent Context Engineering Framework
by: Yuksel, Kamer Ali
Published: (2025)
by: Yuksel, Kamer Ali
Published: (2025)
Agents' Room: Narrative Generation through Multi-step Collaboration
by: Huot, Fantine, et al.
Published: (2024)
by: Huot, Fantine, et al.
Published: (2024)
AgentRec: Agent Recommendation Using Sentence Embeddings Aligned to Human Feedback
by: Park, Joshua, et al.
Published: (2025)
by: Park, Joshua, et al.
Published: (2025)
Debate, Deliberate, Decide (D3): A Cost-Aware Adversarial Framework for Reliable and Interpretable LLM Evaluation
by: Harrasse, Abir, et al.
Published: (2024)
by: Harrasse, Abir, et al.
Published: (2024)
Rethinking the Value of Multi-Agent Workflow: A Strong Single Agent Baseline
by: Xu, Jiawei, et al.
Published: (2026)
by: Xu, Jiawei, et al.
Published: (2026)
Multi-Agent Computer Use
by: Koh, Jing Yu, et al.
Published: (2026)
by: Koh, Jing Yu, et al.
Published: (2026)
Stochastic Self-Organization in Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2025)
by: Tastan, Nurbek, et al.
Published: (2025)
MLZero: A Multi-Agent System for End-to-end Machine Learning Automation
by: Fang, Haoyang, et al.
Published: (2025)
by: Fang, Haoyang, et al.
Published: (2025)
ReDAct: Uncertainty-Aware Deferral for LLM Agents
by: Piatrashyn, Dzianis, et al.
Published: (2026)
by: Piatrashyn, Dzianis, et al.
Published: (2026)
ReaGAN: Node-as-Agent-Reasoning Graph Agentic Network
by: Guo, Minghao, et al.
Published: (2025)
by: Guo, Minghao, et al.
Published: (2025)
G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems
by: Zhang, Guibin, et al.
Published: (2025)
by: Zhang, Guibin, et al.
Published: (2025)
Response-Conditioned Parallel-to-Sequential Orchestration for Multi-Agent Systems
by: Tastan, Nurbek, et al.
Published: (2026)
by: Tastan, Nurbek, et al.
Published: (2026)
LatentMem: Customizing Latent Memory for Multi-Agent Systems
by: Fu, Muxin, et al.
Published: (2026)
by: Fu, Muxin, et al.
Published: (2026)
RunAgent: Interpreting Natural-Language Plans with Constraint-Guided Execution
by: Srivastava, Arunabh, et al.
Published: (2026)
by: Srivastava, Arunabh, et al.
Published: (2026)
Towards Efficient LLM Grounding for Embodied Multi-Agent Collaboration
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Dive into the Agent Matrix: A Realistic Evaluation of Self-Replication Risk in LLM Agents
by: Zhang, Boxuan, et al.
Published: (2025)
by: Zhang, Boxuan, et al.
Published: (2025)
Multi Agent based Medical Assistant for Edge Devices
by: Gawade, Sakharam, et al.
Published: (2025)
by: Gawade, Sakharam, et al.
Published: (2025)
SPIO: Ensemble and Selective Strategies via LLM-Based Multi-Agent Planning in Automated Data Science
by: Seo, Wonduk, et al.
Published: (2025)
by: Seo, Wonduk, et al.
Published: (2025)
BRIDGE: Bootstrapping Text to Control Time-Series Generation via Multi-Agent Iterative Optimization and Diffusion Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
Build Your Personalized Research Group: A Multiagent Framework for Continual and Interactive Science Automation
by: Li, Ed, et al.
Published: (2025)
by: Li, Ed, et al.
Published: (2025)
Can Agents Judge Systematic Reviews Like Humans? Evaluating SLRs with LLM-based Multi-Agent System
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
Symphony: A Decentralized Multi-Agent Framework for Scalable Collective Intelligence
by: Wang, Ji, et al.
Published: (2025)
by: Wang, Ji, et al.
Published: (2025)
Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integration and Evaluation using Novel Metrics and Dataset
by: Gabriel, Adrian Garret, et al.
Published: (2024)
by: Gabriel, Adrian Garret, et al.
Published: (2024)
MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
by: Lin, Huawei, et al.
Published: (2026)
by: Lin, Huawei, et al.
Published: (2026)
Harnessing Multi-Agent LLMs for Complex Engineering Problem-Solving: A Framework for Senior Design Projects
by: Mushtaq, Abdullah, et al.
Published: (2025)
by: Mushtaq, Abdullah, et al.
Published: (2025)
AutoAct: Automatic Agent Learning from Scratch for QA via Self-Planning
by: Qiao, Shuofei, et al.
Published: (2024)
by: Qiao, Shuofei, et al.
Published: (2024)
Agentic Medical Knowledge Graphs Enhance Medical Question Answering: Bridging the Gap Between LLMs and Evolving Medical Knowledge
by: Rezaei, Mohammad Reza, et al.
Published: (2025)
by: Rezaei, Mohammad Reza, et al.
Published: (2025)
Prescriptive Agents based on RAG for Automated Maintenance (PARAM)
by: Harbola, Chitranshu, et al.
Published: (2025)
by: Harbola, Chitranshu, et al.
Published: (2025)
MAEBE: Multi-Agent Emergent Behavior Framework
by: Erisken, Sinem, et al.
Published: (2025)
by: Erisken, Sinem, et al.
Published: (2025)
LLM Agents Making Agent Tools
by: Wölflein, Georg, et al.
Published: (2025)
by: Wölflein, Georg, et al.
Published: (2025)
ConfAgents: A Conformal-Guided Multi-Agent Framework for Cost-Efficient Medical Diagnosis
by: Zhao, Huiya, et al.
Published: (2025)
by: Zhao, Huiya, et al.
Published: (2025)
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
by: Sun, Yu, et al.
Published: (2025)
by: Sun, Yu, et al.
Published: (2025)
Persuade Me if You Can: A Framework for Evaluating Persuasion Effectiveness and Susceptibility Among Large Language Models
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
by: Bozdag, Nimet Beyza, et al.
Published: (2025)
TABX: A High-Throughput Sandbox Battle Simulator for Multi-Agent Reinforcement Learning
by: Lee, Hayeong, et al.
Published: (2026)
by: Lee, Hayeong, et al.
Published: (2026)
Similar Items
-
Rethinking Test-Time Scaling for Medical AI: Model and Task-Aware Strategies for LLMs and VLMs
by: Oh, Gyutaek, et al.
Published: (2025) -
AutoML-Agent: A Multi-Agent LLM Framework for Full-Pipeline AutoML
by: Trirat, Patara, et al.
Published: (2024) -
Doctorina MedBench: End-to-End Evaluation of Agent-Based Medical AI
by: Kozlova, Anna, et al.
Published: (2026) -
MAATS: A Multi-Agent Automated Translation System Based on MQM Evaluation
by: Wang, George, et al.
Published: (2025) -
MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical Tasks
by: Zhu, Yinghao, et al.
Published: (2025)