Multi-Iteration Multi-Stage Fine-Tuning of Transformers for Sound Event Detection with Heterogeneous Datasets
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Schmid, Florian, Primus, Paul, Morocutti, Tobias, Greif, Jonathan, Widmer, Gerhard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
On Temporal Guidance and Iterative Refinement in Audio Source Separation
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
Effective Pre-Training of Audio Transformers for Sound Event Detection
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Exploring Performance-Complexity Trade-Offs in Sound Event Detection Models
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
von: Greif, Jonathan, et al.
Veröffentlicht: (2024)
von: Greif, Jonathan, et al.
Veröffentlicht: (2024)
Multi-Stage Music Source Restoration with BandSplit-RoFormer Separation and HiFi++ GAN
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2026)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
von: Primus, Paul, et al.
Veröffentlicht: (2025)
von: Primus, Paul, et al.
Veröffentlicht: (2025)
Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025)
Device-Robust Acoustic Scene Classification via Impulse Response Augmentation
von: Morocutti, Tobias, et al.
Veröffentlicht: (2023)
von: Morocutti, Tobias, et al.
Veröffentlicht: (2023)
Fusing Audio and Metadata Embeddings Improves Language-based Audio Retrieval
von: Primus, Paul, et al.
Veröffentlicht: (2024)
von: Primus, Paul, et al.
Veröffentlicht: (2024)
Low-Complexity Acoustic Scene Classification with Device Information in the DCASE 2025 Challenge
von: Schmid, Florian, et al.
Veröffentlicht: (2025)
von: Schmid, Florian, et al.
Veröffentlicht: (2025)
Data-Efficient Low-Complexity Acoustic Scene Classification in the DCASE 2024 Challenge
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
von: Schmid, Florian, et al.
Veröffentlicht: (2024)
Sound Event Detection with Boundary-Aware Optimization and Inference
von: Schmid, Florian, et al.
Veröffentlicht: (2026)
von: Schmid, Florian, et al.
Veröffentlicht: (2026)
MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation
von: Hong, Hengyi, et al.
Veröffentlicht: (2024)
von: Hong, Hengyi, et al.
Veröffentlicht: (2024)
DG-SED: Domain Generalization for Sound Event Detection with Heterogeneous Training Data
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
Pushing the Limit of Sound Event Detection with Multi-Dilated Frequency Dynamic Convolution
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2024)
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2024)
FMSG-JLESS Submission for DCASE 2024 Task4 on Sound Event Detection with Heterogeneous Training Dataset and Potentially Missing Labels
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
Fine-Grained Engine Fault Sound Event Detection Using Multimodal Signals
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
von: Fedorishin, Dennis, et al.
Veröffentlicht: (2024)
CST-former: Multidimensional Attention-based Transformer for Sound Event Localization and Detection in Real Scenes
von: Shul, Yusun, et al.
Veröffentlicht: (2025)
von: Shul, Yusun, et al.
Veröffentlicht: (2025)
MTDA-HSED: Mutual-Assistance Tuning and Dual-Branch Aggregating for Heterogeneous Sound Event Detection
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
von: Wang, Zehao, et al.
Veröffentlicht: (2024)
DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
Learning Multi-Target TDOA Features for Sound Event Localization and Detection
von: Berg, Axel, et al.
Veröffentlicht: (2024)
von: Berg, Axel, et al.
Veröffentlicht: (2024)
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
von: Cai, Pengfei, et al.
Veröffentlicht: (2025)
Class-Incremental Learning for Sound Event Localization and Detection
von: Pandey, Ruchi, et al.
Veröffentlicht: (2024)
von: Pandey, Ruchi, et al.
Veröffentlicht: (2024)
LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
von: Koga, Naoki, et al.
Veröffentlicht: (2024)
von: Koga, Naoki, et al.
Veröffentlicht: (2024)
Sound Event Bounding Boxes
von: Ebbers, Janek, et al.
Veröffentlicht: (2024)
von: Ebbers, Janek, et al.
Veröffentlicht: (2024)
WildDESED: An LLM-Powered Dataset for Wild Domestic Environment Sound Event Detection System
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
von: Xiao, Yang, et al.
Veröffentlicht: (2024)
Frequency Dynamic Convolutions for Sound Event Detection
von: Nam, Hyeonuk
Veröffentlicht: (2025)
von: Nam, Hyeonuk
Veröffentlicht: (2025)
Sounding Out Reconstruction Error-Based Evaluation of Generative Models of Expressive Performance
von: Peter, Silvan David, et al.
Veröffentlicht: (2023)
von: Peter, Silvan David, et al.
Veröffentlicht: (2023)
Mind the Domain Gap: a Systematic Analysis on Bioacoustic Sound Event Detection
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
von: Liang, Jinhua, et al.
Veröffentlicht: (2024)
PSELDNets: Pre-trained Neural Networks on a Large-scale Synthetic Dataset for Sound Event Localization and Detection
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
von: Hu, Jinbo, et al.
Veröffentlicht: (2024)
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
von: Bibbó, Gabriel, et al.
Veröffentlicht: (2024)
Zero- and Few-shot Sound Event Localization and Detection
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
von: Shimada, Kazuki, et al.
Veröffentlicht: (2023)
Towards Understanding of Frequency Dependence on Sound Event Detection
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
von: Nam, Hyeonuk, et al.
Veröffentlicht: (2025)
Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
von: Chen, Yuanjian, et al.
Veröffentlicht: (2025)
von: Chen, Yuanjian, et al.
Veröffentlicht: (2025)
Pseudo Strong Labels from Frame-Level Predictions for Weakly Supervised Sound Event Detection
von: Zhang, Yuliang, et al.
Veröffentlicht: (2025)
von: Zhang, Yuliang, et al.
Veröffentlicht: (2025)
Evaluating Sound Similarity Metrics for Differentiable, Iterative Sound-Matching
von: Salimi, Amir, et al.
Veröffentlicht: (2025)
von: Salimi, Amir, et al.
Veröffentlicht: (2025)
Text-Queried Target Sound Event Localization
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
von: Zhao, Jinzheng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
von: Schmid, Florian, et al.
Veröffentlicht: (2024) -
On Temporal Guidance and Iterative Refinement in Audio Source Separation
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025) -
Effective Pre-Training of Audio Transformers for Sound Event Detection
von: Schmid, Florian, et al.
Veröffentlicht: (2024) -
Exploring Performance-Complexity Trade-Offs in Sound Event Detection Models
von: Morocutti, Tobias, et al.
Veröffentlicht: (2025) -
Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
von: Greif, Jonathan, et al.
Veröffentlicht: (2024)