Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866916707454418944 |
|---|---|
| author | Niizumi, Daisuke Takeuchi, Daiki Yasuda, Masahiro Nguyen, Binh Thien Ohishi, Yasunori Harada, Noboru |
| author_facet | Niizumi, Daisuke Takeuchi, Daiki Yasuda, Masahiro Nguyen, Binh Thien Ohishi, Yasunori Harada, Noboru |
| contents | Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound models, which have demonstrated effectiveness as off-the-shelf feature extractors. However, their evaluation benchmarking has been limited, resulting in incompatibility with state-of-the-art (SOTA) performance, thus hindering proof of their effectiveness. This study investigates the practical effectiveness of off-the-shelf audio foundation models by comparing their performance across four respiratory and heart sound tasks with SOTA fine-tuning results. Experiments show that models struggled on two tasks with noisy data but achieved SOTA performance on the other tasks with clean data. Moreover, general-purpose audio models outperformed a respiratory sound model, highlighting their broader applicability. With gained insights and the released code, we contribute to future research on developing and leveraging foundation models for respiratory and heart sounds. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2504_18004 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis Niizumi, Daisuke Takeuchi, Daiki Yasuda, Masahiro Nguyen, Binh Thien Ohishi, Yasunori Harada, Noboru Audio and Speech Processing Sound 68T07 J.3 Pre-trained deep learning models, known as foundation models, have become essential building blocks in machine learning domains such as natural language processing and image domains. This trend has extended to respiratory and heart sound models, which have demonstrated effectiveness as off-the-shelf feature extractors. However, their evaluation benchmarking has been limited, resulting in incompatibility with state-of-the-art (SOTA) performance, thus hindering proof of their effectiveness. This study investigates the practical effectiveness of off-the-shelf audio foundation models by comparing their performance across four respiratory and heart sound tasks with SOTA fine-tuning results. Experiments show that models struggled on two tasks with noisy data but achieved SOTA performance on the other tasks with clean data. Moreover, general-purpose audio models outperformed a respiratory sound model, highlighting their broader applicability. With gained insights and the released code, we contribute to future research on developing and leveraging foundation models for respiratory and heart sounds. |
| title | Assessing the Utility of Audio Foundation Models for Heart and Respiratory Sound Analysis |
| topic | Audio and Speech Processing Sound 68T07 J.3 |
| url | https://arxiv.org/abs/2504.18004 |