RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuwei, Xia, Tong, Saeed, Aaqib, Mascolo, Cecilia
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912062529077248
author Zhang, Yuwei
Xia, Tong
Saeed, Aaqib
Mascolo, Cecilia
author_facet Zhang, Yuwei
Xia, Tong
Saeed, Aaqib
Mascolo, Cecilia
contents The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area. However, the data involved, spanning demographics, medical history, symptoms, and respiratory audio, are heterogeneous and complex. Existing approaches are insufficient and lack generalizability, as they typically rely on limited training data, basic fusion techniques, and task-specific models. In this paper, we propose RespLLM, a novel multimodal large language model (LLM) framework that unifies text and audio representations for respiratory health prediction. RespLLM leverages the extensive prior knowledge of pretrained LLMs and enables effective audio-text fusion through cross-modal attentions. Instruction tuning is employed to integrate diverse data from multiple sources, ensuring generalizability and versatility of the model. Experiments on five real-world datasets demonstrate that RespLLM outperforms leading baselines by an average of 4.6% on trained tasks, 7.9% on unseen datasets, and facilitates zero-shot predictions for new tasks. Our work lays the foundation for multimodal models that can perceive, listen to, and understand heterogeneous data, paving the way for scalable respiratory health diagnosis.
format Preprint
id arxiv_https___arxiv_org_abs_2410_05361
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
Zhang, Yuwei
Xia, Tong
Saeed, Aaqib
Mascolo, Cecilia
Machine Learning
Artificial Intelligence
Sound
Audio and Speech Processing
The high incidence and mortality rates associated with respiratory diseases underscores the importance of early screening. Machine learning models can automate clinical consultations and auscultation, offering vital support in this area. However, the data involved, spanning demographics, medical history, symptoms, and respiratory audio, are heterogeneous and complex. Existing approaches are insufficient and lack generalizability, as they typically rely on limited training data, basic fusion techniques, and task-specific models. In this paper, we propose RespLLM, a novel multimodal large language model (LLM) framework that unifies text and audio representations for respiratory health prediction. RespLLM leverages the extensive prior knowledge of pretrained LLMs and enables effective audio-text fusion through cross-modal attentions. Instruction tuning is employed to integrate diverse data from multiple sources, ensuring generalizability and versatility of the model. Experiments on five real-world datasets demonstrate that RespLLM outperforms leading baselines by an average of 4.6% on trained tasks, 7.9% on unseen datasets, and facilitates zero-shot predictions for new tasks. Our work lays the foundation for multimodal models that can perceive, listen to, and understand heterogeneous data, paving the way for scalable respiratory health diagnosis.
title RespLLM: Unifying Audio and Text with Multimodal LLMs for Generalized Respiratory Health Prediction
topic Machine Learning
Artificial Intelligence
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2410.05361