Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Feng, Tiantian, Xu, Anfeng, Shi, Xuan, Bishop, Somer, Narayanan, Shrikanth
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912406957981696
author Feng, Tiantian
Xu, Anfeng
Shi, Xuan
Bishop, Somer
Narayanan, Shrikanth
author_facet Feng, Tiantian
Xu, Anfeng
Shi, Xuan
Bishop, Somer
Narayanan, Shrikanth
contents Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by challenges in social communication, repetitive behavior, and sensory processing. One important research area in ASD is evaluating children's behavioral changes over time during treatment. The standard protocol with this objective is BOSCC, which involves dyadic interactions between a child and clinicians performing a pre-defined set of activities. A fundamental aspect of understanding children's behavior in these interactions is automatic speech understanding, particularly identifying who speaks and when. Conventional approaches in this area heavily rely on speech samples recorded from a spectator perspective, and there is limited research on egocentric speech modeling. In this study, we design an experiment to perform speech sampling in BOSCC interviews from an egocentric perspective using wearable sensors and explore pre-training Ego4D speech samples to enhance child-adult speaker classification in dyadic interactions. Our findings highlight the potential of egocentric speech collection and pre-training to improve speaker classification accuracy.
format Preprint
id arxiv_https___arxiv_org_abs_2409_09340
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
Feng, Tiantian
Xu, Anfeng
Shi, Xuan
Bishop, Somer
Narayanan, Shrikanth
Sound
Artificial Intelligence
Audio and Speech Processing
Autism spectrum disorder (ASD) is a neurodevelopmental condition characterized by challenges in social communication, repetitive behavior, and sensory processing. One important research area in ASD is evaluating children's behavioral changes over time during treatment. The standard protocol with this objective is BOSCC, which involves dyadic interactions between a child and clinicians performing a pre-defined set of activities. A fundamental aspect of understanding children's behavior in these interactions is automatic speech understanding, particularly identifying who speaks and when. Conventional approaches in this area heavily rely on speech samples recorded from a spectator perspective, and there is limited research on egocentric speech modeling. In this study, we design an experiment to perform speech sampling in BOSCC interviews from an egocentric perspective using wearable sensors and explore pre-training Ego4D speech samples to enhance child-adult speaker classification in dyadic interactions. Our findings highlight the potential of egocentric speech collection and pre-training to improve speaker classification accuracy.
title Egocentric Speaker Classification in Child-Adult Dyadic Interactions: From Sensing to Computational Modeling
topic Sound
Artificial Intelligence
Audio and Speech Processing
url https://arxiv.org/abs/2409.09340