Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Gao, Yifan, Guo, Long, Liu, Hong
Format: Preprint
Veröffentlicht: 2024
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866910833898946560
author Gao, Yifan
Guo, Long
Liu, Hong
author_facet Gao, Yifan
Guo, Long
Liu, Hong
contents Cognitive impairment detection through spontaneous speech is a promising avenue for early diagnosis of Alzheimer's disease (AD) and mild cognitive impairment (MCI), where timely intervention can significantly improve patient outcomes. The PROCESS Grand Challenge at ICASSP 2025 addresses these tasks by promoting innovative classification and regression methods for detecting cognitive decline. In this paper, we propose a multimodal fusion strategy that combines interpretable linguistic features with temporal embeddings extracted from pre-trained models. Our approach achieves an F1-score of 0.649 for the classification task (predicting healthy, MCI, dementia) and an RMSE of 2.628 for the regression task (MMSE score prediction), securing the top overall ranking in the competition.
format Preprint
id arxiv_https___arxiv_org_abs_2412_09928
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification
Gao, Yifan
Guo, Long
Liu, Hong
Sound
Audio and Speech Processing
Cognitive impairment detection through spontaneous speech is a promising avenue for early diagnosis of Alzheimer's disease (AD) and mild cognitive impairment (MCI), where timely intervention can significantly improve patient outcomes. The PROCESS Grand Challenge at ICASSP 2025 addresses these tasks by promoting innovative classification and regression methods for detecting cognitive decline. In this paper, we propose a multimodal fusion strategy that combines interpretable linguistic features with temporal embeddings extracted from pre-trained models. Our approach achieves an F1-score of 0.649 for the classification task (predicting healthy, MCI, dementia) and an RMSE of 2.628 for the regression task (MMSE score prediction), securing the top overall ranking in the competition.
title Leveraging Multimodal Methods and Spontaneous Speech for Alzheimer's Disease Identification
topic Sound
Audio and Speech Processing
url https://arxiv.org/abs/2412.09928