Saved in:
Bibliographic Details
Main Authors: Luong, Justin, Xue, Hao, Salim, Flora D.
Format: Preprint
Published: 2025
Subjects:
Online Access:https://arxiv.org/abs/2508.03764
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909724544335872
author Luong, Justin
Xue, Hao
Salim, Flora D.
author_facet Luong, Justin
Xue, Hao
Salim, Flora D.
contents Physicians routinely assess respiratory sounds during the diagnostic process, providing insight into the condition of a patient's airways. In recent years, AI-based diagnostic systems operating on respiratory sounds, have demonstrated success in respiratory disease detection. These systems represent a crucial advancement in early and accessible diagnosis which is essential for timely treatment. However, label and data scarcity remain key challenges, especially for conditions beyond COVID-19, limiting diagnostic performance and reliable evaluation. In this paper, we propose CoughViT, a novel pre-training framework for learning general-purpose cough sound representations, to enhance diagnostic performance in tasks with limited data. To address label scarcity, we employ masked data modelling to train a feature encoder in a self-supervised learning manner. We evaluate our approach against other pre-training strategies on three diagnostically important cough classification tasks. Experimental results show that our representations match or exceed current state-of-the-art supervised audio representations in enhancing performance on downstream tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2508_03764
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
Luong, Justin
Xue, Hao
Salim, Flora D.
Sound
Artificial Intelligence
Audio and Speech Processing
Physicians routinely assess respiratory sounds during the diagnostic process, providing insight into the condition of a patient's airways. In recent years, AI-based diagnostic systems operating on respiratory sounds, have demonstrated success in respiratory disease detection. These systems represent a crucial advancement in early and accessible diagnosis which is essential for timely treatment. However, label and data scarcity remain key challenges, especially for conditions beyond COVID-19, limiting diagnostic performance and reliable evaluation. In this paper, we propose CoughViT, a novel pre-training framework for learning general-purpose cough sound representations, to enhance diagnostic performance in tasks with limited data. To address label scarcity, we employ masked data modelling to train a feature encoder in a self-supervised learning manner. We evaluate our approach against other pre-training strategies on three diagnostically important cough classification tasks. Experimental results show that our representations match or exceed current state-of-the-art supervised audio representations in enhancing performance on downstream tasks.
title CoughViT: A Self-Supervised Vision Transformer for Cough Audio Representation Learning
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2508.03764