Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866914046968594432 |
|---|---|
| author | Niu, Xinlei Ma, Jianbo Harper-Harris, Dylan Zhang, Xiangyu Martin, Charles Patrick Zhang, Jing |
| author_facet | Niu, Xinlei Ma, Jianbo Harper-Harris, Dylan Zhang, Xiangyu Martin, Charles Patrick Zhang, Jing |
| contents | The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on Foley sound generation, they struggle to produce intelligible speech. Meanwhile, current environmental speech synthesis approaches remain text-driven and fail to temporally align with dynamic video content. In this paper, we propose Beyond Video-to-SFX (BVS), a method to generate synchronized audio with environmentally aware intelligible speech for given videos. We introduce a two-stage modeling method: (1) stage one is a video-guided audio semantic (V2AS) model to predict unified audio semantic tokens conditioned on phonetic cues; (2) stage two is a video-conditioned semantic-to-acoustic (VS2A) model that refines semantic tokens into detailed acoustic tokens. Experiments demonstrate the effectiveness of BVS in scenarios such as video-to-context-aware speech synthesis and immersive audio background conversion, with ablation studies further validating our design. Our demonstration is available at~\href{https://xinleiniu.github.io/BVS-demo/}{BVS-Demo}. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2509_15492 |
| institution | arXiv |
| publishDate | 2025 |
| record_format | arxiv |
| spellingShingle | Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech Niu, Xinlei Ma, Jianbo Harper-Harris, Dylan Zhang, Xiangyu Martin, Charles Patrick Zhang, Jing Sound Multimedia Audio and Speech Processing The generation of realistic, context-aware audio is important in real-world applications such as video game development. While existing video-to-audio (V2A) methods mainly focus on Foley sound generation, they struggle to produce intelligible speech. Meanwhile, current environmental speech synthesis approaches remain text-driven and fail to temporally align with dynamic video content. In this paper, we propose Beyond Video-to-SFX (BVS), a method to generate synchronized audio with environmentally aware intelligible speech for given videos. We introduce a two-stage modeling method: (1) stage one is a video-guided audio semantic (V2AS) model to predict unified audio semantic tokens conditioned on phonetic cues; (2) stage two is a video-conditioned semantic-to-acoustic (VS2A) model that refines semantic tokens into detailed acoustic tokens. Experiments demonstrate the effectiveness of BVS in scenarios such as video-to-context-aware speech synthesis and immersive audio background conversion, with ablation studies further validating our design. Our demonstration is available at~\href{https://xinleiniu.github.io/BVS-demo/}{BVS-Demo}. |
| title | Beyond Video-to-SFX: Video to Audio Synthesis with Environmentally Aware Speech |
| topic | Sound Multimedia Audio and Speech Processing |
| url | https://arxiv.org/abs/2509.15492 |