A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Wang, Pingjie, Zhao, Liudan, Zhao, Zihan, He, Miao, Sun, Xin, Zhang, Ya, Sun, Kun, Wang, Yanfeng, Wang, Yu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866908282261602304
author Wang, Pingjie
Zhao, Liudan
Zhao, Zihan
He, Miao
Sun, Xin
Zhang, Ya
Sun, Kun
Wang, Yanfeng
Wang, Yu
author_facet Wang, Pingjie
Zhao, Liudan
Zhao, Zihan
He, Miao
Sun, Xin
Zhang, Ya
Sun, Kun
Wang, Yanfeng
Wang, Yu
contents Accurate and efficient auscultation-based diagnostics are vital for early disease detection, especially in resource-limited settings where specialized clinical expertise is scarce. Traditional auscultation, which heavily depends on clinician experience, suffers from significant inter-observer variability, while existing AI models often falter due to the limitations of non-representative training data. In this study, we introduce AuscultaBase, a novel AI-driven diagnostic framework that harnesses self-supervised and contrastive learning techniques alongside large-scale, multi-source data integration to advance body sound analysis. By generating robust feature representations, AuscultaBase markedly enhances performance in abnormality detection, disease classification, and activity recognition tasks. Comprehensive evaluations on our newly established benchmark, AuscultaBench, demonstrate that AuscultaBase consistently outperforms state-of-the-art methods across key performance metrics, underscoring its potential as a scalable and cost-effective tool for clinical screening and early disease intervention. The code and model checkpoint has been released in https://github.com/applewpj/AuscultaBase.
format Preprint
id arxiv_https___arxiv_org_abs_2411_07547
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
Wang, Pingjie
Zhao, Liudan
Zhao, Zihan
He, Miao
Sun, Xin
Zhang, Ya
Sun, Kun
Wang, Yanfeng
Wang, Yu
Sound
Audio and Speech Processing
Accurate and efficient auscultation-based diagnostics are vital for early disease detection, especially in resource-limited settings where specialized clinical expertise is scarce. Traditional auscultation, which heavily depends on clinician experience, suffers from significant inter-observer variability, while existing AI models often falter due to the limitations of non-representative training data. In this study, we introduce AuscultaBase, a novel AI-driven diagnostic framework that harnesses self-supervised and contrastive learning techniques alongside large-scale, multi-source data integration to advance body sound analysis. By generating robust feature representations, AuscultaBase markedly enhances performance in abnormality detection, disease classification, and activity recognition tasks. Comprehensive evaluations on our newly established benchmark, AuscultaBench, demonstrate that AuscultaBase consistently outperforms state-of-the-art methods across key performance metrics, underscoring its potential as a scalable and cost-effective tool for clinical screening and early disease intervention. The code and model checkpoint has been released in https://github.com/applewpj/AuscultaBase.
title A Generalist Audio Foundation Model for Comprehensive Body Sound Auscultation
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2411.07547