GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wu, Fan, Zhao, Kaicheng, Fleisch, Elgar, Barata, Filipe
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911089930797056
author Wu, Fan
Zhao, Kaicheng
Fleisch, Elgar
Barata, Filipe
author_facet Wu, Fan
Zhao, Kaicheng
Fleisch, Elgar
Barata, Filipe
contents AI-based voice analysis shows promise for disease diagnostics, but existing classifiers often fail to accurately identify specific pathologies because of gender-related acoustic variations and the scarcity of data for rare diseases. We propose a novel two-stage framework that first identifies gender-specific pathological patterns using ResNet-50 on Mel spectrograms, then performs gender-conditioned disease classification. We address class imbalance through multi-scale resampling and time warping augmentation. Evaluated on a merged dataset from four public repositories, our two-stage architecture with time warping achieves state-of-the-art performance (97.63\% accuracy, 95.25\% MCC), with a 5\% MCC improvement over single-stage baseline. This work advances voice pathology classification while reducing gender bias through hierarchical modeling of vocal characteristics.
format Preprint
id arxiv_https___arxiv_org_abs_2508_01172
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification
Wu, Fan
Zhao, Kaicheng
Fleisch, Elgar
Barata, Filipe
Sound
Artificial Intelligence
Audio and Speech Processing
AI-based voice analysis shows promise for disease diagnostics, but existing classifiers often fail to accurately identify specific pathologies because of gender-related acoustic variations and the scarcity of data for rare diseases. We propose a novel two-stage framework that first identifies gender-specific pathological patterns using ResNet-50 on Mel spectrograms, then performs gender-conditioned disease classification. We address class imbalance through multi-scale resampling and time warping augmentation. Evaluated on a merged dataset from four public repositories, our two-stage architecture with time warping achieves state-of-the-art performance (97.63\% accuracy, 95.25\% MCC), with a 5\% MCC improvement over single-stage baseline. This work advances voice pathology classification while reducing gender bias through hierarchical modeling of vocal characteristics.
title GeHirNet: A Gender-Aware Hierarchical Model for Voice Pathology Classification
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2508.01172