Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Labrak, Yanis, Grünert, David, Baroudi, Séverin, Chun, Jiyun, Cyrta, Pawel, Burdisso, Sergio, Hassoon, Ahmed, Liu, David, Rothschild, Adam, Van Deusen, Reed, Motlicek, Petr, Perrault, Andrew, Marxer, Ricard, Schaaf, Thomas
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908944402743296
author Labrak, Yanis
Grünert, David
Baroudi, Séverin
Chun, Jiyun
Cyrta, Pawel
Burdisso, Sergio
Hassoon, Ahmed
Liu, David
Rothschild, Adam
Van Deusen, Reed
Motlicek, Petr
Perrault, Andrew
Marxer, Ricard
Schaaf, Thomas
author_facet Labrak, Yanis
Grünert, David
Baroudi, Séverin
Chun, Jiyun
Cyrta, Pawel
Burdisso, Sergio
Hassoon, Ahmed
Liu, David
Rothschild, Adam
Van Deusen, Reed
Motlicek, Petr
Perrault, Andrew
Marxer, Ricard
Schaaf, Thomas
contents Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for automatic evaluation. We propose a synthetic data generation pipeline designed to serve both as a training resource and as a controlled evaluation environment, and instantiate it for first-visit doctor-patient conversations with SOAP note generation as the task. The pipeline has three stages, persona-driven dialogue generation, multi-speaker audio synthesis with overlap/pause modeling, room acoustics, and sound events, and LLM-based reference SOAP note production, built entirely on open-weight models. We release 8,800 synthetic conversations with 1.3k hours of corresponding audio and reference notes. Evaluating current open-weight systems, we find that cascaded approaches still substantially outperform end-to-end models.
format Preprint
id arxiv_https___arxiv_org_abs_2604_06138
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization
Labrak, Yanis
Grünert, David
Baroudi, Séverin
Chun, Jiyun
Cyrta, Pawel
Burdisso, Sergio
Hassoon, Ahmed
Liu, David
Rothschild, Adam
Van Deusen, Reed
Motlicek, Petr
Perrault, Andrew
Marxer, Ricard
Schaaf, Thomas
Sound
Artificial Intelligence
Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for automatic evaluation. We propose a synthetic data generation pipeline designed to serve both as a training resource and as a controlled evaluation environment, and instantiate it for first-visit doctor-patient conversations with SOAP note generation as the task. The pipeline has three stages, persona-driven dialogue generation, multi-speaker audio synthesis with overlap/pause modeling, room acoustics, and sound events, and LLM-based reference SOAP note production, built entirely on open-weight models. We release 8,800 synthetic conversations with 1.3k hours of corresponding audio and reference notes. Evaluating current open-weight systems, we find that cascaded approaches still substantially outperform end-to-end models.
title Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization
topic Sound
Artificial Intelligence
url https://arxiv.org/abs/2604.06138