PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection

Fuente: arXiv
Enregistré dans:
Détails bibliographiques
Auteurs principaux: Pahar, Madhurananda, Illingworth, Caitlin H., Mirheidari, Bahman, Elghazaly, Hend, Peters, Fritz, Young, Sophie, Leung, Wing-Zin, Kaur, Labhpreet, Blackburn, Daniel, Christensen, Heidi
Format: Preprint
Publié: 2026
Sujets:
Accès en ligne:
Tags: Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
_version_ 1866918501711609856
author Pahar, Madhurananda
Illingworth, Caitlin H.
Mirheidari, Bahman
Elghazaly, Hend
Peters, Fritz
Young, Sophie
Leung, Wing-Zin
Kaur, Labhpreet
Blackburn, Daniel
Christensen, Heidi
author_facet Pahar, Madhurananda
Illingworth, Caitlin H.
Mirheidari, Bahman
Elghazaly, Hend
Peters, Fritz
Young, Sophie
Leung, Wing-Zin
Kaur, Labhpreet
Blackburn, Daniel
Christensen, Heidi
contents Speech-based analysis offers a scalable and non-invasive approach for detecting cognitive decline, yet progress has been constrained by the limited availability of clinically validated datasets collected under realistic conditions. We introduce PROCESS-2, a large-scale speech dataset designed to support research on automatic assessment of cognitive impairment from spontaneous and task-oriented speech. The dataset comprises recordings from 200 healthy controls, 150 mild cognitive impairment, and 50 dementia diagnoses collected using the CognoMemory digital assessment platform. Each participant completed a single assessment session, including picture description and verbal fluency tasks, accompanied by manually verified transcripts and participant-level metadata. PROCESS-2 contains approximately 21 hours of speech audio with predefined train/test partitions. Comprehensive technical validation evaluated demographic balance, clinical consistency, recording stability, embedding-space structure, and reproducible baseline modelling performance, demonstrating clinically meaningful group separation and stable performance across modelling approaches while preserving real-world conversational variability. PROCESS-2 is released under controlled access via Hugging Face to enable responsible reuse while protecting participant privacy, providing a reproducible benchmark resource for speech-based cognitive assessment research.
format Preprint
id arxiv_https___arxiv_org_abs_2605_14888
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection
Pahar, Madhurananda
Illingworth, Caitlin H.
Mirheidari, Bahman
Elghazaly, Hend
Peters, Fritz
Young, Sophie
Leung, Wing-Zin
Kaur, Labhpreet
Blackburn, Daniel
Christensen, Heidi
Sound
Machine Learning
Speech-based analysis offers a scalable and non-invasive approach for detecting cognitive decline, yet progress has been constrained by the limited availability of clinically validated datasets collected under realistic conditions. We introduce PROCESS-2, a large-scale speech dataset designed to support research on automatic assessment of cognitive impairment from spontaneous and task-oriented speech. The dataset comprises recordings from 200 healthy controls, 150 mild cognitive impairment, and 50 dementia diagnoses collected using the CognoMemory digital assessment platform. Each participant completed a single assessment session, including picture description and verbal fluency tasks, accompanied by manually verified transcripts and participant-level metadata. PROCESS-2 contains approximately 21 hours of speech audio with predefined train/test partitions. Comprehensive technical validation evaluated demographic balance, clinical consistency, recording stability, embedding-space structure, and reproducible baseline modelling performance, demonstrating clinically meaningful group separation and stable performance across modelling approaches while preserving real-world conversational variability. PROCESS-2 is released under controlled access via Hugging Face to enable responsible reuse while protecting participant privacy, providing a reproducible benchmark resource for speech-based cognitive assessment research.
title PROCESS-2: A Benchmark Speech Corpus for Early Cognitive Impairment Detection
topic Sound
Machine Learning
url https://arxiv.org/abs/2605.14888