Teaching Machines to Speak Using Articulatory Control

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Anand, Akshay, Guo, Chenxu, Cho, Cheol Jun, Lian, Jiachen, Anumanchipalli, Gopala
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866918155592400896
author Anand, Akshay
Guo, Chenxu
Cho, Cheol Jun
Lian, Jiachen
Anumanchipalli, Gopala
author_facet Anand, Akshay
Guo, Chenxu
Cho, Cheol Jun
Lian, Jiachen
Anumanchipalli, Gopala
contents Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing a new framework: speech generation through explicit articulatory control. This reframes speech as a motor control task similar to robotic manipulation. Our approach uses reinforcement learning to train a policy that directly controls the movements of vocal tract articulators, such as the tongue, lips, and jaw, to produce syllable-level speech. Specifically, we employ the Proximal Policy Optimization algorithm to learn optimal articulatory movements based on acoustic feedback provided by our audio perceiver, Sylber. The resulting articulatory trajectories are decoded into audio using SPARC, a pre-trained articulatory-to-speech decoder. We train this framework on six target syllables, and it demonstrates successful convergence, with similarity scores between the policy-generated audio and the target syllables exceeding 0.85. Accurate human transcription of the audio for syllables such as "please", "loot", and "cat" demonstrates the intelligibility of this framework.
format Preprint
id arxiv_https___arxiv_org_abs_2510_05619
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Teaching Machines to Speak Using Articulatory Control
Anand, Akshay
Guo, Chenxu
Cho, Cheol Jun
Lian, Jiachen
Anumanchipalli, Gopala
Audio and Speech Processing
Current speech production systems predominantly rely on large transformer models that operate as black boxes, providing little interpretability or grounding in the physical mechanisms of human speech. We address this limitation by proposing a new framework: speech generation through explicit articulatory control. This reframes speech as a motor control task similar to robotic manipulation. Our approach uses reinforcement learning to train a policy that directly controls the movements of vocal tract articulators, such as the tongue, lips, and jaw, to produce syllable-level speech. Specifically, we employ the Proximal Policy Optimization algorithm to learn optimal articulatory movements based on acoustic feedback provided by our audio perceiver, Sylber. The resulting articulatory trajectories are decoded into audio using SPARC, a pre-trained articulatory-to-speech decoder. We train this framework on six target syllables, and it demonstrates successful convergence, with similarity scores between the policy-generated audio and the target syllables exceeding 0.85. Accurate human transcription of the audio for syllables such as "please", "loot", and "cat" demonstrates the intelligibility of this framework.
title Teaching Machines to Speak Using Articulatory Control
topic Audio and Speech Processing
url https://arxiv.org/abs/2510.05619