MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Pengfei, Song, Yan, Li, Kang, Song, Haoyu, McLoughlin, Ian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
von: Cai, Pengfei, et al.
Veröffentlicht: (2024)
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
Effective Pre-Training of Audio Transformers for Sound Event Detection
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
von: Zhong, Guirui, et al.
Veröffentlicht: (2025)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
von: Comunità, Marco, et al.
Veröffentlicht: (2024)
Masked Audio Modeling with CLAP and Multi-Objective Learning
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
von: Xin, Yifei, et al.
Veröffentlicht: (2024)
PSELDNets: Pre-trained Neural Networks on a Large-scale Synthetic Dataset for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
FlexSED: Towards Open-Vocabulary Sound Event Detection
von: Hai, Jiarui, et al.
Veröffentlicht: (2025)
von: Hai, Jiarui, et al.
Veröffentlicht: (2025)
JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
von: Xi, Yu, et al.
Veröffentlicht: (2025)
von: Xi, Yu, et al.
Veröffentlicht: (2025)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
Contrastive Loss Based Frame-wise Feature disentanglement for Polyphonic Sound Event Detection
von: Guan, Yadong, et al.
Veröffentlicht: (2024)
von: Guan, Yadong, et al.
Veröffentlicht: (2024)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
von: Zhao, Junqi, et al.
Veröffentlicht: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
von: Xie, Zeyu, et al.
Veröffentlicht: (2023)
Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
von: Wang, Xiaopeng, et al.
Veröffentlicht: (2024)
von: Wang, Xiaopeng, et al.
Veröffentlicht: (2024)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
von: Yin, Han, et al.
Veröffentlicht: (2024)
von: Yin, Han, et al.
Veröffentlicht: (2024)
w2v-SELD: A Sound Event Localization and Detection Framework for Self-Supervised Spatial Audio Pre-Training
von: Santos, Orlem Lima dos, et al.
Veröffentlicht: (2023)
von: Santos, Orlem Lima dos, et al.
Veröffentlicht: (2023)
Disentangling Dual-Encoder Masked Autoencoder for Respiratory Sound Classification
von: Wei, Peidong, et al.
Veröffentlicht: (2025)
von: Wei, Peidong, et al.
Veröffentlicht: (2025)
AudioSpa: Spatializing Sound Events with Text
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
von: Feng, Linfeng, et al.
Veröffentlicht: (2025)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
von: Roman, Adrian S., et al.
Veröffentlicht: (2025)
The Rhythm In Anything: Audio-Prompted Drums Generation with Masked Language Modeling
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
von: O'Reilly, Patrick, et al.
Veröffentlicht: (2025)
SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
von: Yao, Shengshi, et al.
Veröffentlicht: (2025)
von: Yao, Shengshi, et al.
Veröffentlicht: (2025)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
Analysing the Masked predictive coding training criterion for pre-training a Speech Representation Model
von: Yadav, Hemant, et al.
Veröffentlicht: (2023)
von: Yadav, Hemant, et al.
Veröffentlicht: (2023)
Unified Audio Event Detection
von: Jiang, Yidi, et al.
Veröffentlicht: (2024)
von: Jiang, Yidi, et al.
Veröffentlicht: (2024)
IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2025)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2025)
Sound Event Detection with Boundary-Aware Optimization and Inference
von: Schmid, Florian, et al.
Veröffentlicht: (2026)
von: Schmid, Florian, et al.
Veröffentlicht: (2026)
SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion Transformer
von: Wang, Helin, et al.
Veröffentlicht: (2024)
von: Wang, Helin, et al.
Veröffentlicht: (2024)
Pre-training Autoencoder for Acoustic Event Classification via Blinky
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Liu, Xiaoyang, et al.
Veröffentlicht: (2025)
BLAT: Bootstrapping Language-Audio Pre-training based on AudioSet Tag-guided Synthetic Data
von: Xu, Xuenan, et al.
Veröffentlicht: (2023)
von: Xu, Xuenan, et al.
Veröffentlicht: (2023)
OMAR-RQ: Open Music Audio Representation Model Trained with Multi-Feature Masked Token Prediction
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
von: Alonso-Jiménez, Pablo, et al.
Veröffentlicht: (2025)
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
Frequency Dynamic Convolutions for Sound Event Detection
von: Nam, Hyeonuk
Veröffentlicht: (2025)
von: Nam, Hyeonuk
Veröffentlicht: (2025)
SONAR: Self-Distilled Continual Pre-training for Domain Adaptive Audio Representation
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
von: Xu, Xuenan, et al.
Veröffentlicht: (2024)
Zero- and Few-shot Sound Event Localization and Detection
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
Towards Understanding of Frequency Dependence on Sound Event Detection
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
von: Chen, Yuanjian, et al.
Veröffentlicht: (2025)
von: Chen, Yuanjian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
von: Cai, Pengfei, et al.
Veröffentlicht: (2024) -
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
von: Cai, Pengfei, et al.
Veröffentlicht: (2025) -
Effective Pre-Training of Audio Transformers for Sound Event Detection
von: Schmid, Florian, et al.
Veröffentlicht: (2024) -
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
von: Zhong, Guirui, et al.
Veröffentlicht: (2025) -
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework
von: Niizumi, Daisuke, et al.
Veröffentlicht: (2024)