Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Cai, Pengfei, Song, Yan, Jiang, Nan, Gu, Qing, McLoughlin, Ian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
di: Cai, Pengfei, et al.
Pubblicazione: (2025)
di: Cai, Pengfei, et al.
Pubblicazione: (2025)
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
di: Cai, Pengfei, et al.
Pubblicazione: (2024)
di: Cai, Pengfei, et al.
Pubblicazione: (2024)
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
di: Han, Bing, et al.
Pubblicazione: (2025)
di: Han, Bing, et al.
Pubblicazione: (2025)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
di: Zhao, Junqi, et al.
Pubblicazione: (2024)
di: Zhao, Junqi, et al.
Pubblicazione: (2024)
w2v-SELD: A Sound Event Localization and Detection Framework for Self-Supervised Spatial Audio Pre-Training
di: Santos, Orlem Lima dos, et al.
Pubblicazione: (2023)
di: Santos, Orlem Lima dos, et al.
Pubblicazione: (2023)
JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection
di: Nam, Hyeonuk, et al.
Pubblicazione: (2025)
di: Nam, Hyeonuk, et al.
Pubblicazione: (2025)
Temporal Pooling Strategies for Training-Free Anomalous Sound Detection with Self-Supervised Audio Embeddings
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2026)
di: Wilkinghoff, Kevin, et al.
Pubblicazione: (2026)
Effective Pre-Training of Audio Transformers for Sound Event Detection
di: Schmid, Florian, et al.
Pubblicazione: (2024)
di: Schmid, Florian, et al.
Pubblicazione: (2024)
Enhance Temporal Relations in Audio Captioning with Sound Event Detection
di: Xie, Zeyu, et al.
Pubblicazione: (2023)
di: Xie, Zeyu, et al.
Pubblicazione: (2023)
Exploring Text-Queried Sound Event Detection with Audio Source Separation
di: Yin, Han, et al.
Pubblicazione: (2024)
di: Yin, Han, et al.
Pubblicazione: (2024)
Unified Audio Event Detection
di: Jiang, Yidi, et al.
Pubblicazione: (2024)
di: Jiang, Yidi, et al.
Pubblicazione: (2024)
AudioSpa: Spatializing Sound Events with Text
di: Feng, Linfeng, et al.
Pubblicazione: (2025)
di: Feng, Linfeng, et al.
Pubblicazione: (2025)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
di: Roman, Adrian S., et al.
Pubblicazione: (2025)
di: Roman, Adrian S., et al.
Pubblicazione: (2025)
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
di: Schmid, Florian, et al.
Pubblicazione: (2024)
di: Schmid, Florian, et al.
Pubblicazione: (2024)
Cross-Referencing Self-Training Network for Sound Event Detection in Audio Mixtures
di: Park, Sangwook, et al.
Pubblicazione: (2021)
di: Park, Sangwook, et al.
Pubblicazione: (2021)
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
di: Imoto, Keisuke
Pubblicazione: (2025)
di: Imoto, Keisuke
Pubblicazione: (2025)
Sound Event Detection with Boundary-Aware Optimization and Inference
di: Schmid, Florian, et al.
Pubblicazione: (2026)
di: Schmid, Florian, et al.
Pubblicazione: (2026)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
di: Xi, Yu, et al.
Pubblicazione: (2025)
di: Xi, Yu, et al.
Pubblicazione: (2025)
Utilizing Speaker Profiles for Impersonation Audio Detection
di: Gu, Hao, et al.
Pubblicazione: (2024)
di: Gu, Hao, et al.
Pubblicazione: (2024)
Masked Audio Modeling with CLAP and Multi-Objective Learning
di: Xin, Yifei, et al.
Pubblicazione: (2024)
di: Xin, Yifei, et al.
Pubblicazione: (2024)
Contrastive Loss Based Frame-wise Feature disentanglement for Polyphonic Sound Event Detection
di: Guan, Yadong, et al.
Pubblicazione: (2024)
di: Guan, Yadong, et al.
Pubblicazione: (2024)
Frequency Dynamic Convolutions for Sound Event Detection
di: Nam, Hyeonuk
Pubblicazione: (2025)
di: Nam, Hyeonuk
Pubblicazione: (2025)
Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots
di: Lee, Gyeong-Tae, et al.
Pubblicazione: (2025)
di: Lee, Gyeong-Tae, et al.
Pubblicazione: (2025)
A Detailed Audio-Text Data Simulation Pipeline using Single-Event Sounds
di: Xu, Xuenan, et al.
Pubblicazione: (2024)
di: Xu, Xuenan, et al.
Pubblicazione: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
di: Chen, Yafeng, et al.
Pubblicazione: (2024)
Self-Distillation Prototypes Network: Learning Robust Speaker Representations without Supervision
di: Chen, Yafeng, et al.
Pubblicazione: (2023)
di: Chen, Yafeng, et al.
Pubblicazione: (2023)
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
di: Bibbó, Gabriel, et al.
Pubblicazione: (2024)
di: Bibbó, Gabriel, et al.
Pubblicazione: (2024)
SpecMaskGIT: Masked Generative Modeling of Audio Spectrograms for Efficient Audio Synthesis and Beyond
di: Comunità, Marco, et al.
Pubblicazione: (2024)
di: Comunità, Marco, et al.
Pubblicazione: (2024)
Zero- and Few-shot Sound Event Localization and Detection
di: Shimada, Kazuki, et al.
Pubblicazione: (2023)
di: Shimada, Kazuki, et al.
Pubblicazione: (2023)
Towards Understanding of Frequency Dependence on Sound Event Detection
di: Nam, Hyeonuk, et al.
Pubblicazione: (2025)
di: Nam, Hyeonuk, et al.
Pubblicazione: (2025)
Automatic Sound Event Detection and Classification of Great Ape Calls Using Neural Networks
di: Jiang, Zifan, et al.
Pubblicazione: (2023)
di: Jiang, Zifan, et al.
Pubblicazione: (2023)
Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
di: Chen, Yuanjian, et al.
Pubblicazione: (2025)
di: Chen, Yuanjian, et al.
Pubblicazione: (2025)
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
di: Fujita, Yoto, et al.
Pubblicazione: (2024)
di: Fujita, Yoto, et al.
Pubblicazione: (2024)
Self-Supervised Multi-View Learning for Disentangled Music Audio Representations
di: Wilkins, Julia, et al.
Pubblicazione: (2024)
di: Wilkins, Julia, et al.
Pubblicazione: (2024)
Hierarchical Pooling Structure for Weakly Labeled Sound Event Detection
di: He, Ke-Xin, et al.
Pubblicazione: (2019)
di: He, Ke-Xin, et al.
Pubblicazione: (2019)
Ensemble Confidence Calibration for Sound Event Detection in Open-environment
di: Chen, Yuanjian, et al.
Pubblicazione: (2025)
di: Chen, Yuanjian, et al.
Pubblicazione: (2025)
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
di: Zhong, Guirui, et al.
Pubblicazione: (2025)
di: Zhong, Guirui, et al.
Pubblicazione: (2025)
Genuine-Focused Learning using Mask AutoEncoder for Generalized Fake Audio Detection
di: Wang, Xiaopeng, et al.
Pubblicazione: (2024)
di: Wang, Xiaopeng, et al.
Pubblicazione: (2024)
Binaural Sound Event Localization and Detection Neural Network based on HRTF Localization Cues for Humanoid Robots
di: Lee, Gyeong-Tae
Pubblicazione: (2025)
di: Lee, Gyeong-Tae
Pubblicazione: (2025)
SoundSpring: Loss-Resilient Audio Transceiver with Dual-Functional Masked Language Modeling
di: Yao, Shengshi, et al.
Pubblicazione: (2025)
di: Yao, Shengshi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
di: Cai, Pengfei, et al.
Pubblicazione: (2025) -
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
di: Cai, Pengfei, et al.
Pubblicazione: (2024) -
Exploring Self-Supervised Audio Models for Generalized Anomalous Sound Detection
di: Han, Bing, et al.
Pubblicazione: (2025) -
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
di: Zhao, Junqi, et al.
Pubblicazione: (2024) -
w2v-SELD: A Sound Event Localization and Detection Framework for Self-Supervised Spatial Audio Pre-Training
di: Santos, Orlem Lima dos, et al.
Pubblicazione: (2023)