The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Wang, Yi, Liu, Oli Danyi, Bell, Peter
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866915358451957760
author Wang, Yi
Liu, Oli Danyi
Bell, Peter
author_facet Wang, Yi
Liu, Oli Danyi
Bell, Peter
contents Human speech perception is multimodal. In natural speech, lip movements can precede corresponding voicing by a non-negligible gap of 100-300 ms, especially for specific consonants, affecting the time course of neural phonetic encoding in human listeners. However, it remains unexplored whether self-supervised learning models, which have been used to simulate audio-visual integration in humans, can capture this asynchronicity between audio and visual cues. We compared AV-HuBERT, an audio-visual model, with audio-only HuBERT, by using linear classifiers to track their phonetic decodability over time. We found that phoneme information becomes available in AV-HuBERT embeddings only about 20 ms before HuBERT, likely due to AV-HuBERT's lower temporal resolution and feature concatenation process. It suggests AV-HuBERT does not adequately capture the temporal dynamics of multimodal speech perception, limiting its suitability for modeling the multimodal speech perception process.
format Preprint
id arxiv_https___arxiv_org_abs_2506_20361
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
Wang, Yi
Liu, Oli Danyi
Bell, Peter
Audio and Speech Processing
Sound
Image and Video Processing
Human speech perception is multimodal. In natural speech, lip movements can precede corresponding voicing by a non-negligible gap of 100-300 ms, especially for specific consonants, affecting the time course of neural phonetic encoding in human listeners. However, it remains unexplored whether self-supervised learning models, which have been used to simulate audio-visual integration in humans, can capture this asynchronicity between audio and visual cues. We compared AV-HuBERT, an audio-visual model, with audio-only HuBERT, by using linear classifiers to track their phonetic decodability over time. We found that phoneme information becomes available in AV-HuBERT embeddings only about 20 ms before HuBERT, likely due to AV-HuBERT's lower temporal resolution and feature concatenation process. It suggests AV-HuBERT does not adequately capture the temporal dynamics of multimodal speech perception, limiting its suitability for modeling the multimodal speech perception process.
title The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
topic Audio and Speech Processing
Sound
Image and Video Processing
url https://arxiv.org/abs/2506.20361