Guardado en:
Detalles Bibliográficos
Autores principales: Devahi, Adharsha Sam Edwin Sam, Sangha, Sohail Singh, Priyadarshinee, Prachee, Thilakan, Jithin, Tan, Ivan Fu Xing, Clarke, Christopher Johann, Lon, Sou Ka, T, Balamurali B, Quin, Yow Wei, Jer-Ming, Chen
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:https://arxiv.org/abs/2510.03336
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866918153996468224
author Devahi, Adharsha Sam Edwin Sam
Sangha, Sohail Singh
Priyadarshinee, Prachee
Thilakan, Jithin
Tan, Ivan Fu Xing
Clarke, Christopher Johann
Lon, Sou Ka
T, Balamurali B
Quin, Yow Wei
Jer-Ming, Chen
author_facet Devahi, Adharsha Sam Edwin Sam
Sangha, Sohail Singh
Priyadarshinee, Prachee
Thilakan, Jithin
Tan, Ivan Fu Xing
Clarke, Christopher Johann
Lon, Sou Ka
T, Balamurali B
Quin, Yow Wei
Jer-Ming, Chen
contents Early detection of Alzheimer's Dementia (AD) and Mild Cognitive Impairment (MCI) is critical for timely intervention, yet current diagnostic approaches remain resource-intensive and invasive. Speech, encompassing both acoustic and linguistic dimensions, offers a promising non-invasive biomarker for cognitive decline. In this study, we present a machine learning framework for the PROCESS Challenge, leveraging both audio embeddings and linguistic features derived from spontaneous speech recordings. Audio representations were extracted using Whisper embeddings from the Cookie Theft description task, while linguistic features-spanning pronoun usage, syntactic complexity, filler words, and clause structure-were obtained from transcriptions across Semantic Fluency, Phonemic Fluency, and Cookie Theft picture description. Classification models aimed to distinguish between Healthy Controls (HC), MCI, and AD participants, while regression models predicted Mini-Mental State Examination (MMSE) scores. Results demonstrated that voted ensemble models trained on concatenated linguistic features achieved the best classification performance (F1 = 0.497), while Whisper embedding-based ensemble regressors yielded the lowest MMSE prediction error (RMSE = 2.843). Comparative evaluation within the PROCESS Challenge placed our models among the top submissions in regression task, and mid-range for classification, highlighting the complementary strengths of linguistic and audio embeddings. These findings reinforce the potential of multimodal speech-based approaches for scalable, non-invasive cognitive assessment and underline the importance of integrating task-specific linguistic and acoustic markers in dementia detection.
format Preprint
id arxiv_https___arxiv_org_abs_2510_03336
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge
Devahi, Adharsha Sam Edwin Sam
Sangha, Sohail Singh
Priyadarshinee, Prachee
Thilakan, Jithin
Tan, Ivan Fu Xing
Clarke, Christopher Johann
Lon, Sou Ka
T, Balamurali B
Quin, Yow Wei
Jer-Ming, Chen
Sound
Artificial Intelligence
Machine Learning
Early detection of Alzheimer's Dementia (AD) and Mild Cognitive Impairment (MCI) is critical for timely intervention, yet current diagnostic approaches remain resource-intensive and invasive. Speech, encompassing both acoustic and linguistic dimensions, offers a promising non-invasive biomarker for cognitive decline. In this study, we present a machine learning framework for the PROCESS Challenge, leveraging both audio embeddings and linguistic features derived from spontaneous speech recordings. Audio representations were extracted using Whisper embeddings from the Cookie Theft description task, while linguistic features-spanning pronoun usage, syntactic complexity, filler words, and clause structure-were obtained from transcriptions across Semantic Fluency, Phonemic Fluency, and Cookie Theft picture description. Classification models aimed to distinguish between Healthy Controls (HC), MCI, and AD participants, while regression models predicted Mini-Mental State Examination (MMSE) scores. Results demonstrated that voted ensemble models trained on concatenated linguistic features achieved the best classification performance (F1 = 0.497), while Whisper embedding-based ensemble regressors yielded the lowest MMSE prediction error (RMSE = 2.843). Comparative evaluation within the PROCESS Challenge placed our models among the top submissions in regression task, and mid-range for classification, highlighting the complementary strengths of linguistic and audio embeddings. These findings reinforce the potential of multimodal speech-based approaches for scalable, non-invasive cognitive assessment and underline the importance of integrating task-specific linguistic and acoustic markers in dementia detection.
title Linguistic and Audio Embedding-Based Machine Learning for Alzheimer's Dementia and Mild Cognitive Impairment Detection: Insights from the PROCESS Challenge
topic Sound
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.03336