Towards Pre-training an Effective Respiratory Audio Foundation Model
Fuente:
arXiv
Guardado en:
| Autores principales: | , , , , , |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
| _version_ | 1866915296013451264 |
|---|---|
| author | Niizumi, Daisuke Takeuchi, Daiki Yasuda, Masahiro Nguyen, Binh Thien Ohishi, Yasunori Harada, Noboru |
| author_facet | Niizumi, Daisuke Takeuchi, Daiki Yasuda, Masahiro Nguyen, Binh Thien Ohishi, Yasunori Harada, Noboru |
| contents | Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to datasets that are small-sized and lack diversity has not been sufficiently verified. This study aims to explore better pre-training practices for respiratory sounds by comparing numerous pre-trained audio models. Our investigation reveals that models pre-trained on AudioSet, a general audio dataset, are more effective than the models specifically pre-trained on respiratory sounds. Moreover, combining AudioSet and respiratory sound datasets for further pre-training enhances performance, and preserving the frequency-wise information when aggregating features is vital. Along with more insights found in the experiments, we establish a new state-of-the-art for the OPERA benchmark, contributing to advancing respiratory audio foundation models. Our code is available online at https://github.com/nttcslab/eval-audio-repr/tree/main/plugin/OPERA. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2505_15307 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Towards Pre-training an Effective Respiratory Audio Foundation Model Niizumi, Daisuke Takeuchi, Daiki Yasuda, Masahiro Nguyen, Binh Thien Ohishi, Yasunori Harada, Noboru Audio and Speech Processing Sound 68T07 J.3 Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to datasets that are small-sized and lack diversity has not been sufficiently verified. This study aims to explore better pre-training practices for respiratory sounds by comparing numerous pre-trained audio models. Our investigation reveals that models pre-trained on AudioSet, a general audio dataset, are more effective than the models specifically pre-trained on respiratory sounds. Moreover, combining AudioSet and respiratory sound datasets for further pre-training enhances performance, and preserving the frequency-wise information when aggregating features is vital. Along with more insights found in the experiments, we establish a new state-of-the-art for the OPERA benchmark, contributing to advancing respiratory audio foundation models. Our code is available online at https://github.com/nttcslab/eval-audio-repr/tree/main/plugin/OPERA. |
| title | Towards Pre-training an Effective Respiratory Audio Foundation Model |
| topic | Audio and Speech Processing Sound 68T07 J.3 |
| url | https://arxiv.org/abs/2505.15307 |