Towards Pre-training an Effective Respiratory Audio Foundation Model

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Niizumi, Daisuke, Takeuchi, Daiki, Yasuda, Masahiro, Nguyen, Binh Thien, Ohishi, Yasunori, Harada, Noboru
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915296013451264
author Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Ohishi, Yasunori
Harada, Noboru
author_facet Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Ohishi, Yasunori
Harada, Noboru
contents Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to datasets that are small-sized and lack diversity has not been sufficiently verified. This study aims to explore better pre-training practices for respiratory sounds by comparing numerous pre-trained audio models. Our investigation reveals that models pre-trained on AudioSet, a general audio dataset, are more effective than the models specifically pre-trained on respiratory sounds. Moreover, combining AudioSet and respiratory sound datasets for further pre-training enhances performance, and preserving the frequency-wise information when aggregating features is vital. Along with more insights found in the experiments, we establish a new state-of-the-art for the OPERA benchmark, contributing to advancing respiratory audio foundation models. Our code is available online at https://github.com/nttcslab/eval-audio-repr/tree/main/plugin/OPERA.
format Preprint
id arxiv_https___arxiv_org_abs_2505_15307
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Pre-training an Effective Respiratory Audio Foundation Model
Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Ohishi, Yasunori
Harada, Noboru
Audio and Speech Processing
Sound
68T07
J.3
Recent advancements in foundation models have sparked interest in respiratory audio foundation models. However, the effectiveness of applying conventional pre-training schemes to datasets that are small-sized and lack diversity has not been sufficiently verified. This study aims to explore better pre-training practices for respiratory sounds by comparing numerous pre-trained audio models. Our investigation reveals that models pre-trained on AudioSet, a general audio dataset, are more effective than the models specifically pre-trained on respiratory sounds. Moreover, combining AudioSet and respiratory sound datasets for further pre-training enhances performance, and preserving the frequency-wise information when aggregating features is vital. Along with more insights found in the experiments, we establish a new state-of-the-art for the OPERA benchmark, contributing to advancing respiratory audio foundation models. Our code is available online at https://github.com/nttcslab/eval-audio-repr/tree/main/plugin/OPERA.
title Towards Pre-training an Effective Respiratory Audio Foundation Model
topic Audio and Speech Processing
Sound
68T07
J.3
url https://arxiv.org/abs/2505.15307