Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool Classroom Speech

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Sun, Anchen, Feng, Tiantian, Gutierrez, Gabriela, Londono, Juan J, Xu, Anfeng, Elbaum, Batya, Narayanan, Shrikanth, Perry, Lynn K, Messinger, Daniel S
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866911228724510720
author Sun, Anchen
Feng, Tiantian
Gutierrez, Gabriela
Londono, Juan J
Xu, Anfeng
Elbaum, Batya
Narayanan, Shrikanth
Perry, Lynn K
Messinger, Daniel S
author_facet Sun, Anchen
Feng, Tiantian
Gutierrez, Gabriela
Londono, Juan J
Xu, Anfeng
Elbaum, Batya
Narayanan, Shrikanth
Perry, Lynn K
Messinger, Daniel S
contents This paper introduces an automated framework WSW2.0 for analyzing vocal interactions in preschool classrooms, enhancing both accuracy and scalability through the integration of wav2vec2-based speaker classification and Whisper (large-v2 and large-v3) speech transcription. A total of 235 minutes of audio recordings (160 minutes from 12 children and 75 minutes from 5 teachers), were used to compare system outputs to expert human annotations. WSW2.0 achieves a weighted F1 score of .845, accuracy of .846, and an error-corrected kappa of .672 for speaker classification (child vs. teacher). Transcription quality is moderate to high with word error rates of .119 for teachers and .238 for children. WSW2.0 exhibits relatively high absolute agreement intraclass correlations (ICC) with expert transcriptions for a range of classroom language features. These include teacher and child mean utterance length, lexical diversity, question asking, and responses to questions and other utterances, which show absolute agreement intraclass correlations between .64 and .98. To establish scalability, we apply the framework to an extensive dataset spanning two years and over 1,592 hours of classroom audio recordings, demonstrating the framework's robustness for broad real-world applications. These findings highlight the potential of deep learning and natural language processing techniques to revolutionize educational research by providing accurate measures of key features of preschool classroom speech, ultimately guiding more effective intervention strategies and supporting early childhood language development.
format Preprint
id arxiv_https___arxiv_org_abs_2505_09972
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool Classroom Speech
Sun, Anchen
Feng, Tiantian
Gutierrez, Gabriela
Londono, Juan J
Xu, Anfeng
Elbaum, Batya
Narayanan, Shrikanth
Perry, Lynn K
Messinger, Daniel S
Audio and Speech Processing
Machine Learning
Sound
This paper introduces an automated framework WSW2.0 for analyzing vocal interactions in preschool classrooms, enhancing both accuracy and scalability through the integration of wav2vec2-based speaker classification and Whisper (large-v2 and large-v3) speech transcription. A total of 235 minutes of audio recordings (160 minutes from 12 children and 75 minutes from 5 teachers), were used to compare system outputs to expert human annotations. WSW2.0 achieves a weighted F1 score of .845, accuracy of .846, and an error-corrected kappa of .672 for speaker classification (child vs. teacher). Transcription quality is moderate to high with word error rates of .119 for teachers and .238 for children. WSW2.0 exhibits relatively high absolute agreement intraclass correlations (ICC) with expert transcriptions for a range of classroom language features. These include teacher and child mean utterance length, lexical diversity, question asking, and responses to questions and other utterances, which show absolute agreement intraclass correlations between .64 and .98. To establish scalability, we apply the framework to an extensive dataset spanning two years and over 1,592 hours of classroom audio recordings, demonstrating the framework's robustness for broad real-world applications. These findings highlight the potential of deep learning and natural language processing techniques to revolutionize educational research by providing accurate measures of key features of preschool classroom speech, ultimately guiding more effective intervention strategies and supporting early childhood language development.
title Who Said What WSW 2.0? Enhanced Automated Analysis of Preschool Classroom Speech
topic Audio and Speech Processing
Machine Learning
Sound
url https://arxiv.org/abs/2505.09972