_version_ 1866911067727200256
author Vedadi, Elahe
Barrett, David
Harris, Natalie
Wulczyn, Ellery
Reddy, Shashir
Ruparel, Roma
Schaekermann, Mike
Strother, Tim
Tanno, Ryutaro
Sharma, Yash
Lee, Jihyeon
Hughes, Cían
Slack, Dylan
Palepu, Anil
Freyberg, Jan
Saab, Khaled
Liévin, Valentin
Weng, Wei-Hung
Tu, Tao
Liu, Yun
Tomasev, Nenad
Kulkarni, Kavita
Mahdavi, S. Sara
Guu, Kelvin
Barral, Joëlle
Webster, Dale R.
Manyika, James
Hassidim, Avinatan
Chou, Katherine
Matias, Yossi
Kohli, Pushmeet
Rodman, Adam
Natarajan, Vivek
Karthikesalingam, Alan
Stutz, David
author_facet Vedadi, Elahe
Barrett, David
Harris, Natalie
Wulczyn, Ellery
Reddy, Shashir
Ruparel, Roma
Schaekermann, Mike
Strother, Tim
Tanno, Ryutaro
Sharma, Yash
Lee, Jihyeon
Hughes, Cían
Slack, Dylan
Palepu, Anil
Freyberg, Jan
Saab, Khaled
Liévin, Valentin
Weng, Wei-Hung
Tu, Tao
Liu, Yun
Tomasev, Nenad
Kulkarni, Kavita
Mahdavi, S. Sara
Guu, Kelvin
Barral, Joëlle
Webster, Dale R.
Manyika, James
Hassidim, Avinatan
Chou, Katherine
Matias, Yossi
Kohli, Pushmeet
Rodman, Adam
Natarajan, Vivek
Karthikesalingam, Alan
Stutz, David
contents Recent work has demonstrated the promise of conversational AI systems for diagnostic dialogue. However, real-world assurance of patient safety means that providing individual diagnoses and treatment plans is considered a regulated activity by licensed professionals. Furthermore, physicians commonly oversee other team members in such activities, including nurse practitioners (NPs) or physician assistants/associates (PAs). Inspired by this, we propose a framework for effective, asynchronous oversight of the Articulate Medical Intelligence Explorer (AMIE) AI system. We propose guardrailed-AMIE (g-AMIE), a multi-agent system that performs history taking within guardrails, abstaining from individualized medical advice. Afterwards, g-AMIE conveys assessments to an overseeing primary care physician (PCP) in a clinician cockpit interface. The PCP provides oversight and retains accountability of the clinical decision. This effectively decouples oversight from intake and can thus happen asynchronously. In a randomized, blinded virtual Objective Structured Clinical Examination (OSCE) of text consultations with asynchronous oversight, we compared g-AMIE to NPs/PAs or a group of PCPs under the same guardrails. Across 60 scenarios, g-AMIE outperformed both groups in performing high-quality intake, summarizing cases, and proposing diagnoses and management plans for the overseeing PCP to review. This resulted in higher quality composite decisions. PCP oversight of g-AMIE was also more time-efficient than standalone PCP consultations in prior work. While our study does not replicate existing clinical practices and likely underestimates clinicians' capabilities, our results demonstrate the promise of asynchronous oversight as a feasible paradigm for diagnostic AI systems to operate under expert human oversight for enhancing real-world care.
format Preprint
id arxiv_https___arxiv_org_abs_2507_15743
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards physician-centered oversight of conversational diagnostic AI
Vedadi, Elahe
Barrett, David
Harris, Natalie
Wulczyn, Ellery
Reddy, Shashir
Ruparel, Roma
Schaekermann, Mike
Strother, Tim
Tanno, Ryutaro
Sharma, Yash
Lee, Jihyeon
Hughes, Cían
Slack, Dylan
Palepu, Anil
Freyberg, Jan
Saab, Khaled
Liévin, Valentin
Weng, Wei-Hung
Tu, Tao
Liu, Yun
Tomasev, Nenad
Kulkarni, Kavita
Mahdavi, S. Sara
Guu, Kelvin
Barral, Joëlle
Webster, Dale R.
Manyika, James
Hassidim, Avinatan
Chou, Katherine
Matias, Yossi
Kohli, Pushmeet
Rodman, Adam
Natarajan, Vivek
Karthikesalingam, Alan
Stutz, David
Artificial Intelligence
Computation and Language
Human-Computer Interaction
Machine Learning
Recent work has demonstrated the promise of conversational AI systems for diagnostic dialogue. However, real-world assurance of patient safety means that providing individual diagnoses and treatment plans is considered a regulated activity by licensed professionals. Furthermore, physicians commonly oversee other team members in such activities, including nurse practitioners (NPs) or physician assistants/associates (PAs). Inspired by this, we propose a framework for effective, asynchronous oversight of the Articulate Medical Intelligence Explorer (AMIE) AI system. We propose guardrailed-AMIE (g-AMIE), a multi-agent system that performs history taking within guardrails, abstaining from individualized medical advice. Afterwards, g-AMIE conveys assessments to an overseeing primary care physician (PCP) in a clinician cockpit interface. The PCP provides oversight and retains accountability of the clinical decision. This effectively decouples oversight from intake and can thus happen asynchronously. In a randomized, blinded virtual Objective Structured Clinical Examination (OSCE) of text consultations with asynchronous oversight, we compared g-AMIE to NPs/PAs or a group of PCPs under the same guardrails. Across 60 scenarios, g-AMIE outperformed both groups in performing high-quality intake, summarizing cases, and proposing diagnoses and management plans for the overseeing PCP to review. This resulted in higher quality composite decisions. PCP oversight of g-AMIE was also more time-efficient than standalone PCP consultations in prior work. While our study does not replicate existing clinical practices and likely underestimates clinicians' capabilities, our results demonstrate the promise of asynchronous oversight as a feasible paradigm for diagnostic AI systems to operate under expert human oversight for enhancing real-world care.
title Towards physician-centered oversight of conversational diagnostic AI
topic Artificial Intelligence
Computation and Language
Human-Computer Interaction
Machine Learning
url https://arxiv.org/abs/2507.15743