The Dialogue That Heals: A Comprehensive Evaluation of Doctor Agents' Inquiry Capability
Fuente:
arXiv
Saved in:
| Main Authors: | Gong, Linlu, Wang, Ante, Lai, Yunghwei, Ma, Weizhi, Liu, Yang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty
by: Ren, Jingyi, et al.
Published: (2026)
by: Ren, Jingyi, et al.
Published: (2026)
TheraAgent: Self-Improving Therapeutic Agent for Precise and Comprehensive Treatment Planning
by: Li, Junkai, et al.
Published: (2026)
by: Li, Junkai, et al.
Published: (2026)
Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning
by: Lai, Yunghwei, et al.
Published: (2025)
by: Lai, Yunghwei, et al.
Published: (2025)
Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
by: Wang, Ante, et al.
Published: (2025)
by: Wang, Ante, et al.
Published: (2025)
Beyond Words: Evaluating and Bridging Epistemic Divergence in User-Agent Interaction via Theory of Mind
by: Ruan, Minyuan, et al.
Published: (2026)
by: Ruan, Minyuan, et al.
Published: (2026)
Patient-Zero: Scaling Synthetic Patient Agents to Real-World Distributions without Real Patient Data
by: Lai, Yunghwei, et al.
Published: (2025)
by: Lai, Yunghwei, et al.
Published: (2025)
Towards Transparent RAG: Fostering Evidence Traceability in LLM Generation via Reinforcement Learning
by: Ren, Jingyi, et al.
Published: (2025)
by: Ren, Jingyi, et al.
Published: (2025)
Reasoning or Not? A Comprehensive Evaluation of Reasoning LLMs for Dialogue Summarization
by: Jin, Keyan, et al.
Published: (2025)
by: Jin, Keyan, et al.
Published: (2025)
StoryBench: A Dynamic Benchmark for Evaluating Long-Term Memory with Multi Turns
by: Wan, Luanbo, et al.
Published: (2025)
by: Wan, Luanbo, et al.
Published: (2025)
Controlling Language Difficulty in Dialogues with Linguistic Features
by: Xu, Shuyao, et al.
Published: (2025)
by: Xu, Shuyao, et al.
Published: (2025)
Response Enhanced Semi-supervised Dialogue Query Generation
by: Huang, Jianheng, et al.
Published: (2023)
by: Huang, Jianheng, et al.
Published: (2023)
A Proactive EMR Assistant for Doctor-Patient Dialogue: Streaming ASR, Belief Stabilization, and Preliminary Controlled Evaluation
by: Pan, Zhenhai, et al.
Published: (2026)
by: Pan, Zhenhai, et al.
Published: (2026)
DialogueAgents: A Hybrid Agent-Based Speech Synthesis Framework for Multi-Party Dialogue
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
EducationQ: Evaluating LLMs' Teaching Capabilities Through Multi-Agent Dialogue Framework
by: Shi, Yao, et al.
Published: (2025)
by: Shi, Yao, et al.
Published: (2025)
BioVerge: A Comprehensive Benchmark and Study of Self-Evaluating Agents for Biomedical Hypothesis Generation
by: Yang, Fuyi, et al.
Published: (2025)
by: Yang, Fuyi, et al.
Published: (2025)
RAD-Bench: Evaluating Large Language Models Capabilities in Retrieval Augmented Dialogues
by: Kuo, Tzu-Lin, et al.
Published: (2024)
by: Kuo, Tzu-Lin, et al.
Published: (2024)
Large Language Model based Situational Dialogues for Second Language Learning
by: Xu, Shuyao, et al.
Published: (2024)
by: Xu, Shuyao, et al.
Published: (2024)
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
by: Yan, Ruiqi, et al.
Published: (2025)
by: Yan, Ruiqi, et al.
Published: (2025)
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
by: Li, Ce, et al.
Published: (2025)
by: Li, Ce, et al.
Published: (2025)
A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue Evaluators
by: Zhang, Chen, et al.
Published: (2023)
by: Zhang, Chen, et al.
Published: (2023)
IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data
by: Peng, Bo, et al.
Published: (2025)
by: Peng, Bo, et al.
Published: (2025)
SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
by: Gu, Yiyang, et al.
Published: (2026)
by: Gu, Yiyang, et al.
Published: (2026)
PersuasiveToM: A Benchmark for Evaluating Machine Theory of Mind in Persuasive Dialogues
by: Yu, Fangxu, et al.
Published: (2025)
by: Yu, Fangxu, et al.
Published: (2025)
ProactiveEval: A Unified Evaluation Framework for Proactive Dialogue Agents
by: Liu, Tianjian, et al.
Published: (2025)
by: Liu, Tianjian, et al.
Published: (2025)
Are LLMs Effective Negotiators? Systematic Evaluation of the Multifaceted Capabilities of LLMs in Negotiation Dialogues
by: Kwon, Deuksin, et al.
Published: (2024)
by: Kwon, Deuksin, et al.
Published: (2024)
Evaluating Memory Capability in Continuous Lifelog Scenario
by: Zheng, Jianjie, et al.
Published: (2026)
by: Zheng, Jianjie, et al.
Published: (2026)
BattleAgentBench: A Benchmark for Evaluating Cooperation and Competition Capabilities of Language Models in Multi-Agent Systems
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities
by: Lin, Guan-Ting, et al.
Published: (2025)
by: Lin, Guan-Ting, et al.
Published: (2025)
DocTalk: Scalable Graph-based Dialogue Synthesis for Enhancing LLM Conversational Capabilities
by: Lee, Jing Yang, et al.
Published: (2025)
by: Lee, Jing Yang, et al.
Published: (2025)
Reasoning or Reciting? Exploring the Capabilities and Limitations of Language Models Through Counterfactual Tasks
by: Wu, Zhaofeng, et al.
Published: (2023)
by: Wu, Zhaofeng, et al.
Published: (2023)
"Where does it hurt?" -- Dataset and Study on Physician Intent Trajectories in Doctor Patient Dialogues
by: Röhr, Tom, et al.
Published: (2025)
by: Röhr, Tom, et al.
Published: (2025)
MIND: Towards Immersive Psychological Healing with Multi-agent Inner Dialogue
by: Chen, Yujia, et al.
Published: (2025)
by: Chen, Yujia, et al.
Published: (2025)
CDT: A Comprehensive Capability Framework for Large Language Models Across Cognition, Domain, and Task
by: Mo, Haosi, et al.
Published: (2025)
by: Mo, Haosi, et al.
Published: (2025)
SQLBench: A Comprehensive Evaluation for Text-to-SQL Capabilities of Large Language Models
by: Zhang, Bin, et al.
Published: (2024)
by: Zhang, Bin, et al.
Published: (2024)
MemRouter: Memory-as-Embedding Routing for Long-Term Conversational Agents
by: Hu, Tianyu, et al.
Published: (2026)
by: Hu, Tianyu, et al.
Published: (2026)
Dual Hierarchical Dialogue Policy Learning for Legal Inquisitive Conversational Agents
by: Lin, Xubo, et al.
Published: (2026)
by: Lin, Xubo, et al.
Published: (2026)
MUCAR: Benchmarking Multilingual Cross-Modal Ambiguity Resolution for Multimodal Large Language Models
by: Wang, Xiaolong, et al.
Published: (2025)
by: Wang, Xiaolong, et al.
Published: (2025)
Simulating Classroom Education with LLM-Empowered Agents
by: Zhang, Zheyuan, et al.
Published: (2024)
by: Zhang, Zheyuan, et al.
Published: (2024)
A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models
by: Wang, Jiayin, et al.
Published: (2024)
by: Wang, Jiayin, et al.
Published: (2024)
Mind's Mirror: Distilling Self-Evaluation Capability and Comprehensive Thinking from Large Language Models
by: Liu, Weize, et al.
Published: (2023)
by: Liu, Weize, et al.
Published: (2023)
Similar Items
-
Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty
by: Ren, Jingyi, et al.
Published: (2026) -
TheraAgent: Self-Improving Therapeutic Agent for Precise and Comprehensive Treatment Planning
by: Li, Junkai, et al.
Published: (2026) -
Doctor-R1: Mastering Clinical Inquiry with Experiential Agentic Reinforcement Learning
by: Lai, Yunghwei, et al.
Published: (2025) -
Let the Model Distribute Its Doubt: Confidence Estimation through Verbalized Probability Distribution
by: Wang, Ante, et al.
Published: (2025) -
Beyond Words: Evaluating and Bridging Epistemic Divergence in User-Agent Interaction via Theory of Mind
by: Ruan, Minyuan, et al.
Published: (2026)