Efficient Vocal Source Separation Through Windowed Sink Attention

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Benetatos, Christodoulos, Zang, Yongyi, Leistikow, Randal
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912676716740608
author Benetatos, Christodoulos
Zang, Yongyi
Leistikow, Randal
author_facet Benetatos, Christodoulos
Zang, Yongyi
Leistikow, Randal
contents State-of-the-art vocal separation models like Mel-Band-Roformer rely on full temporal self-attention mechanisms, where each temporal frame interacts with every other frames. This incurs heavy computational costs that scales quadratically with input audio length, motivating chunking and windowing approaches. Through analysis of a pre-trained vocal separation model, we discovered that temporal attention patterns are highly localized. Building on this insight, we replaced full attention with windowed sink attention (WSA) with small temporal attention window and attention sinks. We show empirically that fine-tuning from the original checkpoint recovers 92% of the original SDR performance while reducing FLOPs by 44.5x. We release our code and checkpoints under MIT license at https://github.com/smulelabs/windowed-roformer.
format Preprint
id arxiv_https___arxiv_org_abs_2510_25745
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Vocal Source Separation Through Windowed Sink Attention
Benetatos, Christodoulos
Zang, Yongyi
Leistikow, Randal
Sound
State-of-the-art vocal separation models like Mel-Band-Roformer rely on full temporal self-attention mechanisms, where each temporal frame interacts with every other frames. This incurs heavy computational costs that scales quadratically with input audio length, motivating chunking and windowing approaches. Through analysis of a pre-trained vocal separation model, we discovered that temporal attention patterns are highly localized. Building on this insight, we replaced full attention with windowed sink attention (WSA) with small temporal attention window and attention sinks. We show empirically that fine-tuning from the original checkpoint recovers 92% of the original SDR performance while reducing FLOPs by 44.5x. We release our code and checkpoints under MIT license at https://github.com/smulelabs/windowed-roformer.
title Efficient Vocal Source Separation Through Windowed Sink Attention
topic Sound
url https://arxiv.org/abs/2510.25745