Staff View: :: Library Catalog

Saved in:

Bibliographic Details
Main Authors:	Rutowski, Tomek, Harati, Amir, Shriberg, Elizabeth, Lu, Yang, Chlebek, Piotr, Oliveira, Ricardo
Format:	Preprint
Published:	2024
Subjects:	Computation and Language Sound Audio and Speech Processing
Online Access:	https://arxiv.org/abs/2501.00617
Tags:	Add Tag No Tags, Be the first to tag this record!

_version_	1866929654099607552
author	Rutowski, Tomek Harati, Amir Shriberg, Elizabeth Lu, Yang Chlebek, Piotr Oliveira, Ricardo
author_facet	Rutowski, Tomek Harati, Amir Shriberg, Elizabeth Lu, Yang Chlebek, Piotr Oliveira, Ricardo
contents	Mental health risk prediction is a growing field in the speech community, but many studies are based on small corpora. This study illustrates how variations in test and train set sizes impact performance in a controlled study. Using a corpus of over 65K labeled data points, results from a fully crossed design of different train/test size combinations are provided. Two model types are included: one based on language and the other on speech acoustics. Both use methods current in this domain. An age-mismatched test set was also included. Results show that (1) test sizes below 1K samples gave noisy results, even for larger training set sizes; (2) training set sizes of at least 2K were needed for stable results; (3) NLP and acoustic models behaved similarly with train/test size variations, and (4) the mismatched test set showed the same patterns as the matched test set. Additional factors are discussed, including label priors, model strength and pre-training, unique speakers, and data lengths. While no single study can specify exact size requirements, results demonstrate the need for appropriately sized train and test sets for future studies of mental health risk prediction from speech and language.
format	Preprint
id	arxiv_https___arxiv_org_abs_2501_00617
institution	arXiv
publishDate	2024
record_format	arxiv
spellingShingle	Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language Rutowski, Tomek Harati, Amir Shriberg, Elizabeth Lu, Yang Chlebek, Piotr Oliveira, Ricardo Computation and Language Sound Audio and Speech Processing Mental health risk prediction is a growing field in the speech community, but many studies are based on small corpora. This study illustrates how variations in test and train set sizes impact performance in a controlled study. Using a corpus of over 65K labeled data points, results from a fully crossed design of different train/test size combinations are provided. Two model types are included: one based on language and the other on speech acoustics. Both use methods current in this domain. An age-mismatched test set was also included. Results show that (1) test sizes below 1K samples gave noisy results, even for larger training set sizes; (2) training set sizes of at least 2K were needed for stable results; (3) NLP and acoustic models behaved similarly with train/test size variations, and (4) the mismatched test set showed the same patterns as the matched test set. Additional factors are discussed, including label priors, model strength and pre-training, unique speakers, and data lengths. While no single study can specify exact size requirements, results demonstrate the need for appropriately sized train and test sets for future studies of mental health risk prediction from speech and language.
title	Toward Corpus Size Requirements for Training and Evaluating Depression Risk Models Using Spoken Language
topic	Computation and Language Sound Audio and Speech Processing
url	https://arxiv.org/abs/2501.00617

Similar Items