Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autore principale: Salvi, Giampiero
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909071215427584
author Salvi, Giampiero
author_facet Salvi, Giampiero
contents This paper describes the use of connectionist techniques in phonetic speech recognition with strong latency constraints. The constraints are imposed by the task of deriving the lip movements of a synthetic face in real time from the speech signal, by feeding the phonetic string into an articulatory synthesiser. Particular attention has been paid to analysing the interaction between the time evolution model learnt by the multi-layer perceptrons and the transition model imposed by the Viterbi decoder, in different latency conditions. Two experiments were conducted in which the time dependencies in the language model (LM) were controlled by a parameter. The results show a strong interaction between the three factors involved, namely the neural network topology, the length of time dependencies in the LM and the decoder latency.
format Preprint
id arxiv_https___arxiv_org_abs_2401_06588
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
Salvi, Giampiero
Audio and Speech Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Sound
I.5.0; I.2.7; E.4
This paper describes the use of connectionist techniques in phonetic speech recognition with strong latency constraints. The constraints are imposed by the task of deriving the lip movements of a synthetic face in real time from the speech signal, by feeding the phonetic string into an articulatory synthesiser. Particular attention has been paid to analysing the interaction between the time evolution model learnt by the multi-layer perceptrons and the transition model imposed by the Viterbi decoder, in different latency conditions. Two experiments were conducted in which the time dependencies in the language model (LM) were controlled by a parameter. The results show a strong interaction between the three factors involved, namely the neural network topology, the length of time dependencies in the LM and the decoder latency.
title Dynamic Behaviour of Connectionist Speech Recognition with Strong Latency Constraints
topic Audio and Speech Processing
Artificial Intelligence
Computer Vision and Pattern Recognition
Machine Learning
Sound
I.5.0; I.2.7; E.4
url https://arxiv.org/abs/2401.06588