Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Weixin, Qu, Bowen, Pontell, Matthew, Powell, Maria, Malin, Bradley, Yin, Zhijun
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915761972314112
author Liu, Weixin
Qu, Bowen
Pontell, Matthew
Powell, Maria
Malin, Bradley
Yin, Zhijun
author_facet Liu, Weixin
Qu, Bowen
Pontell, Matthew
Powell, Maria
Malin, Bradley
Yin, Zhijun
contents The human voice is a promising non-invasive digital biomarker, yet deep learning for voice-based health analysis is hindered by data scarcity and domain mismatch, where models pre-trained on general audio fail to capture the subtle pathological features characteristic of clinical voice data. To address these challenges, we investigate domain-adaptive self-supervised learning (SSL) with Masked Autoencoders (MAE) and demonstrate that standard configurations are suboptimal for health-related audio. Using the Bridge2AI-Voice dataset, a multi-institutional collection of pathological voices, we systematically examine three performance-critical factors: reconstruction loss (Mean Absolute Error vs. Mean Squared Error), normalization (patch-wise vs. global), and masking (random vs. content-aware). Our optimized design, which combines Mean Absolute Error (MA-Error) loss, patch-wise normalization, and content-aware masking, achieves a Macro F1 of $0.688 \pm 0.009$ (over 10 fine-tuning runs), outperforming a strong out-of-domain SSL baseline pre-trained on large-scale general audio, which has a Macro F1 of $0.663 \pm 0.011$. The results show that MA-Error loss improves robustness and content-aware masking boosts performance by emphasizing information-rich regions. These findings highlight the importance of component-level optimization in data-constrained medical applications that rely on audio data.
format Preprint
id arxiv_https___arxiv_org_abs_2601_22319
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification
Liu, Weixin
Qu, Bowen
Pontell, Matthew
Powell, Maria
Malin, Bradley
Yin, Zhijun
Audio and Speech Processing
Signal Processing
The human voice is a promising non-invasive digital biomarker, yet deep learning for voice-based health analysis is hindered by data scarcity and domain mismatch, where models pre-trained on general audio fail to capture the subtle pathological features characteristic of clinical voice data. To address these challenges, we investigate domain-adaptive self-supervised learning (SSL) with Masked Autoencoders (MAE) and demonstrate that standard configurations are suboptimal for health-related audio. Using the Bridge2AI-Voice dataset, a multi-institutional collection of pathological voices, we systematically examine three performance-critical factors: reconstruction loss (Mean Absolute Error vs. Mean Squared Error), normalization (patch-wise vs. global), and masking (random vs. content-aware). Our optimized design, which combines Mean Absolute Error (MA-Error) loss, patch-wise normalization, and content-aware masking, achieves a Macro F1 of $0.688 \pm 0.009$ (over 10 fine-tuning runs), outperforming a strong out-of-domain SSL baseline pre-trained on large-scale general audio, which has a Macro F1 of $0.663 \pm 0.011$. The results show that MA-Error loss improves robustness and content-aware masking boosts performance by emphasizing information-rich regions. These findings highlight the importance of component-level optimization in data-constrained medical applications that rely on audio data.
title Optimizing Domain-Adaptive Self-Supervised Learning for Clinical Voice-Based Disease Classification
topic Audio and Speech Processing
Signal Processing
url https://arxiv.org/abs/2601.22319