Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913625772392448 |
|---|---|
| author | Bae, Sangmin Kim, June-Woo Cho, Won-Yang Baek, Hyerim Son, Soyoun Lee, Byungjo Ha, Changwan Tae, Kyongpil Kim, Sungnyun Yun, Se-Young |
| author_facet | Bae, Sangmin Kim, June-Woo Cho, Won-Yang Baek, Hyerim Son, Soyoun Lee, Byungjo Ha, Changwan Tae, Kyongpil Kim, Sungnyun Yun, Se-Young |
| contents | Respiratory sound contains crucial information for the early diagnosis of fatal lung diseases. Since the COVID-19 pandemic, there has been a growing interest in contact-free medical care based on electronic stethoscopes. To this end, cutting-edge deep learning models have been developed to diagnose lung diseases; however, it is still challenging due to the scarcity of medical data. In this study, we demonstrate that the pretrained model on large-scale visual and audio datasets can be generalized to the respiratory sound classification task. In addition, we introduce a straightforward Patch-Mix augmentation, which randomly mixes patches between different samples, with Audio Spectrogram Transformer (AST). We further propose a novel and effective Patch-Mix Contrastive Learning to distinguish the mixed representations in the latent space. Our method achieves state-of-the-art performance on the ICBHI dataset, outperforming the prior leading score by an improvement of 4.08%. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2305_14032 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification Bae, Sangmin Kim, June-Woo Cho, Won-Yang Baek, Hyerim Son, Soyoun Lee, Byungjo Ha, Changwan Tae, Kyongpil Kim, Sungnyun Yun, Se-Young Audio and Speech Processing Machine Learning Sound Respiratory sound contains crucial information for the early diagnosis of fatal lung diseases. Since the COVID-19 pandemic, there has been a growing interest in contact-free medical care based on electronic stethoscopes. To this end, cutting-edge deep learning models have been developed to diagnose lung diseases; however, it is still challenging due to the scarcity of medical data. In this study, we demonstrate that the pretrained model on large-scale visual and audio datasets can be generalized to the respiratory sound classification task. In addition, we introduce a straightforward Patch-Mix augmentation, which randomly mixes patches between different samples, with Audio Spectrogram Transformer (AST). We further propose a novel and effective Patch-Mix Contrastive Learning to distinguish the mixed representations in the latent space. Our method achieves state-of-the-art performance on the ICBHI dataset, outperforming the prior leading score by an improvement of 4.08%. |
| title | Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification |
| topic | Audio and Speech Processing Machine Learning Sound |
| url | https://arxiv.org/abs/2305.14032 |