UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Glazer, Neta, Navon, Aviv, Segal, Yael, Shamsian, Aviv, Segev, Hilit, Buchnick, Asaf, Pirchi, Menachem, Hetz, Gil, Keshet, Joseph
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913937657692160
author Glazer, Neta
Navon, Aviv
Segal, Yael
Shamsian, Aviv
Segev, Hilit
Buchnick, Asaf
Pirchi, Menachem
Hetz, Gil
Keshet, Joseph
author_facet Glazer, Neta
Navon, Aviv
Segal, Yael
Shamsian, Aviv
Segev, Hilit
Buchnick, Asaf
Pirchi, Menachem
Hetz, Gil
Keshet, Joseph
contents Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce UmbraTTS, a flow-matching based TTS model that jointly generates both speech and environmental audio, conditioned on text and acoustic context. Our model allows fine-grained control over background volume and produces diverse, coherent, and context-aware audio scenes. A key challenge is the lack of data with speech and background audio aligned in natural context. To overcome the lack of paired training data, we propose a self-supervised framework that extracts speech, background audio, and transcripts from unannotated recordings. Extensive evaluations demonstrate that UmbraTTS significantly outperformed existing baselines, producing natural, high-quality, environmentally aware audios.
format Preprint
id arxiv_https___arxiv_org_abs_2506_09874
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
Glazer, Neta
Navon, Aviv
Segal, Yael
Shamsian, Aviv
Segev, Hilit
Buchnick, Asaf
Pirchi, Menachem
Hetz, Gil
Keshet, Joseph
Sound
Machine Learning
Audio and Speech Processing
Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce UmbraTTS, a flow-matching based TTS model that jointly generates both speech and environmental audio, conditioned on text and acoustic context. Our model allows fine-grained control over background volume and produces diverse, coherent, and context-aware audio scenes. A key challenge is the lack of data with speech and background audio aligned in natural context. To overcome the lack of paired training data, we propose a self-supervised framework that extracts speech, background audio, and transcripts from unannotated recordings. Extensive evaluations demonstrate that UmbraTTS significantly outperformed existing baselines, producing natural, high-quality, environmentally aware audios.
title UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
topic Sound
Machine Learning
Audio and Speech Processing
url https://arxiv.org/abs/2506.09874