Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niizumi, Daisuke, Takeuchi, Daiki, Yasuda, Masahiro, Nguyen, Binh Thien, Ohishi, Yasunori, Harada, Noboru
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916707454418944
author Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Ohishi, Yasunori
Harada, Noboru
author_facet Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Ohishi, Yasunori
Harada, Noboru
contents Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound models, which have demonstrated effectiveness as off-the-shelf feature extractors. However, their evaluation benchmarking has been limited, resulting in incompatibility with state-of-the-art (SOTA) performance, thus hindering proof of their effectiveness. This study investigates the practical effectiveness of off-the-shelf audio foundation models by comparing their performance across four respiratory and heart sound tasks with SOTA fine-tuning results. Experiments show that models struggled on two tasks with noisy data but achieved SOTA performance on the other tasks with clean data. Moreover, general-purpose audio models outperformed a respiratory sound model, highlighting their broader applicability. With gained insights and the released code, we contribute to future research on developing and leveraging foundation models for respiratory and heart sounds.
format Preprint
id arxiv_https___arxiv_org_abs_2504_18004
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis
Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Ohishi, Yasunori
Harada, Noboru
Audio and Speech Processing
Sound
68T07
J.3
Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound models, which have demonstrated effectiveness as off-the-shelf feature extractors. However, their evaluation benchmarking has been limited, resulting in incompatibility with state-of-the-art (SOTA) performance, thus hindering proof of their effectiveness. This study investigates the practical effectiveness of off-the-shelf audio foundation models by comparing their performance across four respiratory and heart sound tasks with SOTA fine-tuning results. Experiments show that models struggled on two tasks with noisy data but achieved SOTA performance on the other tasks with clean data. Moreover, general-purpose audio models outperformed a respiratory sound model, highlighting their broader applicability. With gained insights and the released code, we contribute to future research on developing and leveraging foundation models for respiratory and heart sounds.
title Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis
topic Audio and Speech Processing
Sound
68T07
J.3
url https://arxiv.org/abs/2504.18004