Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Niizumi, Daisuke, Takeuchi, Daiki, Yasuda, Masahiro, Nguyen, Binh Thien, Harada, Noboru, Ono, Nobutaka
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918407562067968
author Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Harada, Noboru
Ono, Nobutaka
author_facet Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Harada, Noboru
Ono, Nobutaka
contents Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation learning using masked prediction-based self-supervised learning (SSL) on general audio spectrograms. While recent informed masking techniques have attracted attention, we observe that they incur substantial computational overhead. Motivated by this observation, we propose dispersion-weighted masking (DWM), a lightweight masking strategy that leverages the spectral sparsity inherent in the frequency structure of audio content. Our experiments show that inverse block masking, commonly used in recent SSL frameworks, improves audio event understanding performance while introducing a trade-off in generalization. The proposed DWM alleviates these limitations and computational complexity, leading to consistent performance improvements. This work provides practical guidance on masking strategy design for masked prediction-based audio representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2603_23810
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
Niizumi, Daisuke
Takeuchi, Daiki
Yasuda, Masahiro
Nguyen, Binh Thien
Harada, Noboru
Ono, Nobutaka
Audio and Speech Processing
Multimedia
Sound
68T07
I.2.6
Since the introduction of Masked Autoencoders, various improvements to masking techniques have been explored. In this paper, we rethink masking strategies for audio representation learning using masked prediction-based self-supervised learning (SSL) on general audio spectrograms. While recent informed masking techniques have attracted attention, we observe that they incur substantial computational overhead. Motivated by this observation, we propose dispersion-weighted masking (DWM), a lightweight masking strategy that leverages the spectral sparsity inherent in the frequency structure of audio content. Our experiments show that inverse block masking, commonly used in recent SSL frameworks, improves audio event understanding performance while introducing a trade-off in generalization. The proposed DWM alleviates these limitations and computational complexity, leading to consistent performance improvements. This work provides practical guidance on masking strategy design for masked prediction-based audio representation learning.
title Rethinking Masking Strategies for Masked Prediction-based Audio Self-supervised Learning
topic Audio and Speech Processing
Multimedia
Sound
68T07
I.2.6
url https://arxiv.org/abs/2603.23810