Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866908944402743296 |
|---|---|
| author | Labrak, Yanis Grünert, David Baroudi, Séverin Chun, Jiyun Cyrta, Pawel Burdisso, Sergio Hassoon, Ahmed Liu, David Rothschild, Adam Van Deusen, Reed Motlicek, Petr Perrault, Andrew Marxer, Ricard Schaaf, Thomas |
| author_facet | Labrak, Yanis Grünert, David Baroudi, Séverin Chun, Jiyun Cyrta, Pawel Burdisso, Sergio Hassoon, Ahmed Liu, David Rothschild, Adam Van Deusen, Reed Motlicek, Petr Perrault, Andrew Marxer, Ricard Schaaf, Thomas |
| contents | Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for automatic evaluation. We propose a synthetic data generation pipeline designed to serve both as a training resource and as a controlled evaluation environment, and instantiate it for first-visit doctor-patient conversations with SOAP note generation as the task. The pipeline has three stages, persona-driven dialogue generation, multi-speaker audio synthesis with overlap/pause modeling, room acoustics, and sound events, and LLM-based reference SOAP note production, built entirely on open-weight models. We release 8,800 synthetic conversations with 1.3k hours of corresponding audio and reference notes. Evaluating current open-weight systems, we find that cascaded approaches still substantially outperform end-to-end models. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2604_06138 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization Labrak, Yanis Grünert, David Baroudi, Séverin Chun, Jiyun Cyrta, Pawel Burdisso, Sergio Hassoon, Ahmed Liu, David Rothschild, Adam Van Deusen, Reed Motlicek, Petr Perrault, Andrew Marxer, Ricard Schaaf, Thomas Sound Artificial Intelligence Long-context audio reasoning is underserved in both training data and evaluation. Existing benchmarks target short-context tasks, and the open-ended generation tasks most relevant to long-context reasoning pose well-known challenges for automatic evaluation. We propose a synthetic data generation pipeline designed to serve both as a training resource and as a controlled evaluation environment, and instantiate it for first-visit doctor-patient conversations with SOAP note generation as the task. The pipeline has three stages, persona-driven dialogue generation, multi-speaker audio synthesis with overlap/pause modeling, room acoustics, and sound events, and LLM-based reference SOAP note production, built entirely on open-weight models. We release 8,800 synthetic conversations with 1.3k hours of corresponding audio and reference notes. Evaluating current open-weight systems, we find that cascaded approaches still substantially outperform end-to-end models. |
| title | Generating Synthetic Doctor-Patient Conversations for Long-form Audio Summarization |
| topic | Sound Artificial Intelligence |
| url | https://arxiv.org/abs/2604.06138 |