Cinematic Audio Source Separation Using Visual Cues

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Kang, Lee, Suyeon, Senocak, Arda, Chung, Joon Son
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911547987591168
author Zhang, Kang
Lee, Suyeon
Senocak, Arda
Chung, Joon Son
author_facet Zhang, Kang
Lee, Suyeon
Senocak, Arda
Chung, Joon Son
contents Cinematic Audio Source Separation (CASS) aims to decompose mixed film audio into speech, music, and sound effects, enabling applications like dubbing and remastering. Existing CASS approaches are audio-only, overlooking the inherent audio-visual nature of films, where sounds often align with visual cues. We present the first framework for audio-visual CASS (AV-CASS), leveraging visual context to enhance separation quality. Our method formulates CASS as a conditional generative modeling problem using conditional flow matching, enabling multimodal audio source separation. To address the lack of cinematic datasets with isolated sound tracks, we introduce a training data synthesis pipeline that pairs in-the-wild audio and video streams (e.g., facial videos for speech, scene videos for effects) and design a dedicated visual encoder for this dual-stream setup. Trained entirely on synthetic data, our model generalizes effectively to real-world cinematic content and achieves strong performance on synthetic, real-world, and audio-only CASS benchmarks. Code and demo are available at \url{https://cass-flowmatching.github.io}.
format Preprint
id arxiv_https___arxiv_org_abs_2603_26113
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Cinematic Audio Source Separation Using Visual Cues
Zhang, Kang
Lee, Suyeon
Senocak, Arda
Chung, Joon Son
Multimedia
Sound
Audio and Speech Processing
Cinematic Audio Source Separation (CASS) aims to decompose mixed film audio into speech, music, and sound effects, enabling applications like dubbing and remastering. Existing CASS approaches are audio-only, overlooking the inherent audio-visual nature of films, where sounds often align with visual cues. We present the first framework for audio-visual CASS (AV-CASS), leveraging visual context to enhance separation quality. Our method formulates CASS as a conditional generative modeling problem using conditional flow matching, enabling multimodal audio source separation. To address the lack of cinematic datasets with isolated sound tracks, we introduce a training data synthesis pipeline that pairs in-the-wild audio and video streams (e.g., facial videos for speech, scene videos for effects) and design a dedicated visual encoder for this dual-stream setup. Trained entirely on synthetic data, our model generalizes effectively to real-world cinematic content and achieves strong performance on synthetic, real-world, and audio-only CASS benchmarks. Code and demo are available at \url{https://cass-flowmatching.github.io}.
title Cinematic Audio Source Separation Using Visual Cues
topic Multimedia
Sound
Audio and Speech Processing
url https://arxiv.org/abs/2603.26113