MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Chen, Hongzhao, Wang, XiaoYang, Lan, Jing, Ding, Hexiao, Jiang, Yufeng, Yang, MingHui, Xu, DanHui, Luo, Jun, Ng, Nga-Chun, Cheng, Gerald W. Y., Mao, Yunlin, Yoo, Jung Sun
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866909808375889920
author Chen, Hongzhao
Wang, XiaoYang
Lan, Jing
Ding, Hexiao
Jiang, Yufeng
Yang, MingHui
Xu, DanHui
Luo, Jun
Ng, Nga-Chun
Cheng, Gerald W. Y.
Mao, Yunlin
Yoo, Jung Sun
author_facet Chen, Hongzhao
Wang, XiaoYang
Lan, Jing
Ding, Hexiao
Jiang, Yufeng
Yang, MingHui
Xu, DanHui
Luo, Jun
Ng, Nga-Chun
Cheng, Gerald W. Y.
Mao, Yunlin
Yoo, Jung Sun
contents Automatic speech recognition (ASR) in clinical dialogue demands robustness to full-duplex interaction, speaker overlap, and low-latency constraints, yet open benchmarks remain scarce. We present MMedFD, the first real-world Chinese healthcare ASR corpus designed for multi-turn, full-duplex settings. Captured from a deployed AI assistant, the dataset comprises 5,805 annotated sessions with synchronized user and mixed-channel views, RTTM/CTM timing, and role labels. We introduce a model-agnostic pipeline for streaming segmentation, speaker attribution, and dialogue memory, and fine-tune Whisper-small on role-concatenated audio for long-context recognition. ASR evaluation includes WER, CER, and HC-WER, which measures concept-level accuracy across healthcare settings. LLM-generated responses are assessed using rubric-based and pairwise protocols. MMedFD establishes a reproducible framework for benchmarking streaming ASR and end-to-end duplex agents in healthcare deployment. The dataset and related resources are publicly available at https://github.com/Kinetics-JOJO/MMedFD
format Preprint
id arxiv_https___arxiv_org_abs_2509_19817
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
Chen, Hongzhao
Wang, XiaoYang
Lan, Jing
Ding, Hexiao
Jiang, Yufeng
Yang, MingHui
Xu, DanHui
Luo, Jun
Ng, Nga-Chun
Cheng, Gerald W. Y.
Mao, Yunlin
Yoo, Jung Sun
Audio and Speech Processing
Automatic speech recognition (ASR) in clinical dialogue demands robustness to full-duplex interaction, speaker overlap, and low-latency constraints, yet open benchmarks remain scarce. We present MMedFD, the first real-world Chinese healthcare ASR corpus designed for multi-turn, full-duplex settings. Captured from a deployed AI assistant, the dataset comprises 5,805 annotated sessions with synchronized user and mixed-channel views, RTTM/CTM timing, and role labels. We introduce a model-agnostic pipeline for streaming segmentation, speaker attribution, and dialogue memory, and fine-tune Whisper-small on role-concatenated audio for long-context recognition. ASR evaluation includes WER, CER, and HC-WER, which measures concept-level accuracy across healthcare settings. LLM-generated responses are assessed using rubric-based and pairwise protocols. MMedFD establishes a reproducible framework for benchmarking streaming ASR and end-to-end duplex agents in healthcare deployment. The dataset and related resources are publicly available at https://github.com/Kinetics-JOJO/MMedFD
title MMedFD: A Real-world Healthcare Benchmark for Multi-turn Full-Duplex Automatic Speech Recognition
topic Audio and Speech Processing
url https://arxiv.org/abs/2509.19817