Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Bae, Sangmin, Kim, June-Woo, Cho, Won-Yang, Baek, Hyerim, Son, Soyoun, Lee, Byungjo, Ha, Changwan, Tae, Kyongpil, Kim, Sungnyun, Yun, Se-Young
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913625772392448
author Bae, Sangmin
Kim, June-Woo
Cho, Won-Yang
Baek, Hyerim
Son, Soyoun
Lee, Byungjo
Ha, Changwan
Tae, Kyongpil
Kim, Sungnyun
Yun, Se-Young
author_facet Bae, Sangmin
Kim, June-Woo
Cho, Won-Yang
Baek, Hyerim
Son, Soyoun
Lee, Byungjo
Ha, Changwan
Tae, Kyongpil
Kim, Sungnyun
Yun, Se-Young
contents Respiratory sound contains crucial information for the early diagnosis of fatal lung diseases. Since the COVID-19 pandemic, there has been a growing interest in contact-free medical care based on electronic stethoscopes. To this end, cutting-edge deep learning models have been developed to diagnose lung diseases; however, it is still challenging due to the scarcity of medical data. In this study, we demonstrate that the pretrained model on large-scale visual and audio datasets can be generalized to the respiratory sound classification task. In addition, we introduce a straightforward Patch-Mix augmentation, which randomly mixes patches between different samples, with Audio Spectrogram Transformer (AST). We further propose a novel and effective Patch-Mix Contrastive Learning to distinguish the mixed representations in the latent space. Our method achieves state-of-the-art performance on the ICBHI dataset, outperforming the prior leading score by an improvement of 4.08%.
format Preprint
id arxiv_https___arxiv_org_abs_2305_14032
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
Bae, Sangmin
Kim, June-Woo
Cho, Won-Yang
Baek, Hyerim
Son, Soyoun
Lee, Byungjo
Ha, Changwan
Tae, Kyongpil
Kim, Sungnyun
Yun, Se-Young
Audio and Speech Processing
Machine Learning
Sound
Respiratory sound contains crucial information for the early diagnosis of fatal lung diseases. Since the COVID-19 pandemic, there has been a growing interest in contact-free medical care based on electronic stethoscopes. To this end, cutting-edge deep learning models have been developed to diagnose lung diseases; however, it is still challenging due to the scarcity of medical data. In this study, we demonstrate that the pretrained model on large-scale visual and audio datasets can be generalized to the respiratory sound classification task. In addition, we introduce a straightforward Patch-Mix augmentation, which randomly mixes patches between different samples, with Audio Spectrogram Transformer (AST). We further propose a novel and effective Patch-Mix Contrastive Learning to distinguish the mixed representations in the latent space. Our method achieves state-of-the-art performance on the ICBHI dataset, outperforming the prior leading score by an improvement of 4.08%.
title Patch-Mix Contrastive Learning with Audio Spectrogram Transformer on Respiratory Sound Classification
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2305.14032