UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching
Fuente:
arXiv
Saved in:
| Main Authors: | , , , , , , , , |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866913937657692160 |
|---|---|
| author | Glazer, Neta Navon, Aviv Segal, Yael Shamsian, Aviv Segev, Hilit Buchnick, Asaf Pirchi, Menachem Hetz, Gil Keshet, Joseph |
| author_facet | Glazer, Neta Navon, Aviv Segal, Yael Shamsian, Aviv Segev, Hilit Buchnick, Asaf Pirchi, Menachem Hetz, Gil Keshet, Joseph |
| contents | Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce UmbraTTS, a flow-matching based TTS model that jointly generates both speech and environmental audio, conditioned on text and acoustic context. Our model allows fine-grained control over background volume and produces diverse, coherent, and context-aware audio scenes. A key challenge is the lack of data with speech and background audio aligned in natural context. To overcome the lack of paired training data, we propose a self-supervised framework that extracts speech, background audio, and transcripts from unannotated recordings. Extensive evaluations demonstrate that UmbraTTS significantly outperformed existing baselines, producing natural, high-quality, environmentally aware audios. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2506_09874 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Glazer, Neta Navon, Aviv Segal, Yael Shamsian, Aviv Segev, Hilit Buchnick, Asaf Pirchi, Menachem Hetz, Gil Keshet, Joseph Sound Machine Learning Audio and Speech Processing Recent advances in Text-to-Speech (TTS) have enabled highly natural speech synthesis, yet integrating speech with complex background environments remains challenging. We introduce UmbraTTS, a flow-matching based TTS model that jointly generates both speech and environmental audio, conditioned on text and acoustic context. Our model allows fine-grained control over background volume and produces diverse, coherent, and context-aware audio scenes. A key challenge is the lack of data with speech and background audio aligned in natural context. To overcome the lack of paired training data, we propose a self-supervised framework that extracts speech, background audio, and transcripts from unannotated recordings. Extensive evaluations demonstrate that UmbraTTS significantly outperformed existing baselines, producing natural, high-quality, environmentally aware audios. |
| title | UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching |
| topic | Sound Machine Learning Audio and Speech Processing |
| url | https://arxiv.org/abs/2506.09874 |