STAR: Speech-to-Audio Generation via Representation Learning

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Xie, Zeyu, Xu, Xuenan, Li, Yixuan, Wu, Mengyue, Zou, Yuexian
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866908551691108352
author Xie, Zeyu
Xu, Xuenan
Li, Yixuan
Wu, Mengyue
Zou, Yuexian
author_facet Xie, Zeyu
Xu, Xuenan
Li, Yixuan
Wu, Mengyue
Zou, Yuexian
contents This work presents STAR, the first end-to-end speech-to-audio generation framework, designed to enhance efficiency and address error propagation inherent in cascaded systems. Unlike prior approaches relying on text or vision, STAR leverages speech as it constitutes a natural modality for interaction. As an initial step to validate the feasibility of the system, we demonstrate through representation learning experiments that spoken sound event semantics can be effectively extracted from raw speech, capturing both auditory events and scene cues. Leveraging the semantic representations, STAR incorporates a bridge network for representation mapping and a two-stage training strategy to achieve end-to-end synthesis. With a 76.9% reduction in speech processing latency, STAR demonstrates superior generation performance over the cascaded systems. Overall, STAR establishes speech as a direct interaction signal for audio generation, thereby bridging representation learning and multimodal synthesis. Generated samples are available at https://zeyuxie29.github.io/STAR.
format Preprint
id arxiv_https___arxiv_org_abs_2509_17164
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle STAR: Speech-to-Audio Generation via Representation Learning
Xie, Zeyu
Xu, Xuenan
Li, Yixuan
Wu, Mengyue
Zou, Yuexian
Sound
Audio and Speech Processing
68Txx
I.2
This work presents STAR, the first end-to-end speech-to-audio generation framework, designed to enhance efficiency and address error propagation inherent in cascaded systems. Unlike prior approaches relying on text or vision, STAR leverages speech as it constitutes a natural modality for interaction. As an initial step to validate the feasibility of the system, we demonstrate through representation learning experiments that spoken sound event semantics can be effectively extracted from raw speech, capturing both auditory events and scene cues. Leveraging the semantic representations, STAR incorporates a bridge network for representation mapping and a two-stage training strategy to achieve end-to-end synthesis. With a 76.9% reduction in speech processing latency, STAR demonstrates superior generation performance over the cascaded systems. Overall, STAR establishes speech as a direct interaction signal for audio generation, thereby bridging representation learning and multimodal synthesis. Generated samples are available at https://zeyuxie29.github.io/STAR.
title STAR: Speech-to-Audio Generation via Representation Learning
topic Sound
Audio and Speech Processing
68Txx
I.2
url https://arxiv.org/abs/2509.17164