RRP-Voice: A Longitudinal Dataset and Benchmark for Recurrent Respiratory Papillomatosis Detection

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ren, Wenze, Lu, Ke-Han, Chang, Kai-Wei, Feng, Tiantian, Fang, Ching, Liao, Zhi-Chi, Yen, Dao Thi Hai, Wang, Syu-Siang, Tsao, Yu, Wang, Chi-Te, Fang, Shih-Hau
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914621761257472
author Ren, Wenze
Lu, Ke-Han
Chang, Kai-Wei
Feng, Tiantian
Fang, Ching
Liao, Zhi-Chi
Yen, Dao Thi Hai
Wang, Syu-Siang
Tsao, Yu
Wang, Chi-Te
Fang, Shih-Hau
author_facet Ren, Wenze
Lu, Ke-Han
Chang, Kai-Wei
Feng, Tiantian
Fang, Ching
Liao, Zhi-Chi
Yen, Dao Thi Hai
Wang, Syu-Siang
Tsao, Yu
Wang, Chi-Te
Fang, Shih-Hau
contents Deep learning has advanced pathological voice detection rapidly, yet rare laryngeal diseases remain underexplored due to data scarcity. Recurrent Respiratory Papillomatosis (RRP) exemplifies this gap: an HPV-induced disease of the larynx in which patients oscillate between recurrence and post-surgical remission over the years. RRP demands continuous voice monitoring that existing cross-sectional corpora cannot support. We introduce the first longitudinal voice dataset for RRP, comprising recordings from 26 patients with up to ten years of follow-up. Each session pairs sustained vowels with sentence-level utterances, which are annotated by otolaryngologists and confirmed synchronously with laryngoscopy. Building on this resource, we establish a systematic benchmark spanning handcrafted features, end-to-end deep networks, self-supervised pretrained models, and recent audio large language models, all evaluated under session-level cross-validation with patient-level audit. Per-subject longitudinal analyses further confirm that the cross-sectional discriminative signal reflects laryngoscopic disease state rather than stable speaker attributes. This work lays a foundation for rare longitudinal pathological voice tasks in low-resource clinical settings.
format Preprint
id arxiv_https___arxiv_org_abs_2606_01639
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle RRP-Voice: A Longitudinal Dataset and Benchmark for Recurrent Respiratory Papillomatosis Detection
Ren, Wenze
Lu, Ke-Han
Chang, Kai-Wei
Feng, Tiantian
Fang, Ching
Liao, Zhi-Chi
Yen, Dao Thi Hai
Wang, Syu-Siang
Tsao, Yu
Wang, Chi-Te
Fang, Shih-Hau
Audio and Speech Processing
Deep learning has advanced pathological voice detection rapidly, yet rare laryngeal diseases remain underexplored due to data scarcity. Recurrent Respiratory Papillomatosis (RRP) exemplifies this gap: an HPV-induced disease of the larynx in which patients oscillate between recurrence and post-surgical remission over the years. RRP demands continuous voice monitoring that existing cross-sectional corpora cannot support. We introduce the first longitudinal voice dataset for RRP, comprising recordings from 26 patients with up to ten years of follow-up. Each session pairs sustained vowels with sentence-level utterances, which are annotated by otolaryngologists and confirmed synchronously with laryngoscopy. Building on this resource, we establish a systematic benchmark spanning handcrafted features, end-to-end deep networks, self-supervised pretrained models, and recent audio large language models, all evaluated under session-level cross-validation with patient-level audit. Per-subject longitudinal analyses further confirm that the cross-sectional discriminative signal reflects laryngoscopic disease state rather than stable speaker attributes. This work lays a foundation for rare longitudinal pathological voice tasks in low-resource clinical settings.
title RRP-Voice: A Longitudinal Dataset and Benchmark for Recurrent Respiratory Papillomatosis Detection
topic Audio and Speech Processing
url https://arxiv.org/abs/2606.01639