Structured-Noise Masked Modeling for Video, Audio and Beyond

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Bhowmik, Aritra, Thoker, Fida Mohammad, Hinojosa, Carlos, Ghanem, Bernard, Snoek, Cees G. M.
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866916658189172736
author Bhowmik, Aritra
Thoker, Fida Mohammad
Hinojosa, Carlos
Ghanem, Bernard
Snoek, Cees G. M.
author_facet Bhowmik, Aritra
Thoker, Fida Mohammad
Hinojosa, Carlos
Ghanem, Bernard
Snoek, Cees G. M.
contents Masked modeling has emerged as a powerful self-supervised learning framework, but existing methods largely rely on random masking, disregarding the structural properties of different modalities. In this work, we introduce structured noise-based masking, a simple yet effective approach that naturally aligns with the spatial, temporal, and spectral characteristics of video and audio data. By filtering white noise into distinct color noise distributions, we generate structured masks that preserve modality-specific patterns without requiring handcrafted heuristics or access to the data. Our approach improves the performance of masked video and audio modeling frameworks without any computational overhead. Extensive experiments demonstrate that structured noise masking achieves consistent improvement over random masking for standard and advanced masked modeling methods, highlighting the importance of modality-aware masking strategies for representation learning.
format Preprint
id arxiv_https___arxiv_org_abs_2503_16311
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Structured-Noise Masked Modeling for Video, Audio and Beyond
Bhowmik, Aritra
Thoker, Fida Mohammad
Hinojosa, Carlos
Ghanem, Bernard
Snoek, Cees G. M.
Machine Learning
Artificial Intelligence
Sound
Masked modeling has emerged as a powerful self-supervised learning framework, but existing methods largely rely on random masking, disregarding the structural properties of different modalities. In this work, we introduce structured noise-based masking, a simple yet effective approach that naturally aligns with the spatial, temporal, and spectral characteristics of video and audio data. By filtering white noise into distinct color noise distributions, we generate structured masks that preserve modality-specific patterns without requiring handcrafted heuristics or access to the data. Our approach improves the performance of masked video and audio modeling frameworks without any computational overhead. Extensive experiments demonstrate that structured noise masking achieves consistent improvement over random masking for standard and advanced masked modeling methods, highlighting the importance of modality-aware masking strategies for representation learning.
title Structured-Noise Masked Modeling for Video, Audio and Beyond
topic Machine Learning
Artificial Intelligence
Sound
url https://arxiv.org/abs/2503.16311