Speechless: Speech Instruction Training Without Speech for Low Resource Languages

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Dao, Alan, Vu, Dinh Bach, Ha, Huy Hoang, Anh, Tuan Le Duc, Gopal, Shreyas, Yeo, Yue Heng, Low, Warren Keng Hoong, Chng, Eng Siong, Yip, Jia Qi
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866909750517563392
author Dao, Alan
Vu, Dinh Bach
Ha, Huy Hoang
Anh, Tuan Le Duc
Gopal, Shreyas
Yeo, Yue Heng
Low, Warren Keng Hoong
Chng, Eng Siong
Yip, Jia Qi
author_facet Dao, Alan
Vu, Dinh Bach
Ha, Huy Hoang
Anh, Tuan Le Duc
Gopal, Shreyas
Yeo, Yue Heng
Low, Warren Keng Hoong
Chng, Eng Siong
Yip, Jia Qi
contents The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of speech recognition data, there is a notable scarcity of speech instruction data, which is essential for fine-tuning models to understand and execute spoken commands. Generating high-quality synthetic speech requires a good text-to-speech (TTS) model, which may not be available to low resource languages. Our novel approach addresses this challenge by halting synthesis at the semantic representation level, bypassing the need for TTS. We achieve this by aligning synthetic semantic representations with the pre-trained Whisper encoder, enabling an LLM to be fine-tuned on text instructions while maintaining the ability to understand spoken instructions during inference. This simplified training process is a promising approach to building voice assistant for low-resource languages.
format Preprint
id arxiv_https___arxiv_org_abs_2505_17417
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Speechless: Speech Instruction Training Without Speech for Low Resource Languages
Dao, Alan
Vu, Dinh Bach
Ha, Huy Hoang
Anh, Tuan Le Duc
Gopal, Shreyas
Yeo, Yue Heng
Low, Warren Keng Hoong
Chng, Eng Siong
Yip, Jia Qi
Audio and Speech Processing
Computation and Language
Sound
The rapid growth of voice assistants powered by large language models (LLM) has highlighted a need for speech instruction data to train these systems. Despite the abundance of speech recognition data, there is a notable scarcity of speech instruction data, which is essential for fine-tuning models to understand and execute spoken commands. Generating high-quality synthetic speech requires a good text-to-speech (TTS) model, which may not be available to low resource languages. Our novel approach addresses this challenge by halting synthesis at the semantic representation level, bypassing the need for TTS. We achieve this by aligning synthetic semantic representations with the pre-trained Whisper encoder, enabling an LLM to be fine-tuned on text instructions while maintaining the ability to understand spoken instructions during inference. This simplified training process is a promising approach to building voice assistant for low-resource languages.
title Speechless: Speech Instruction Training Without Speech for Low Resource Languages
topic Audio and Speech Processing
Computation and Language
Sound
url https://arxiv.org/abs/2505.17417