Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Ronchini, Francesca, Serizel, Romain |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A benchmark of state-of-the-art sound event detection systems evaluated on synthetic soundscapes
di: Ronchini, Francesca, et al.
Pubblicazione: (2022)
di: Ronchini, Francesca, et al.
Pubblicazione: (2022)
The impact of non-target events in synthetic soundscapes for sound event detection
di: Ronchini, Francesca, et al.
Pubblicazione: (2021)
di: Ronchini, Francesca, et al.
Pubblicazione: (2021)
Description and analysis of novelties introduced in DCASE Task 4 2022 on the baseline system
di: Ronchini, Francesca, et al.
Pubblicazione: (2022)
di: Ronchini, Francesca, et al.
Pubblicazione: (2022)
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
di: Ronchini, Francesca, et al.
Pubblicazione: (2020)
di: Ronchini, Francesca, et al.
Pubblicazione: (2020)
Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
di: Passoni, Riccardo, et al.
Pubblicazione: (2025)
di: Passoni, Riccardo, et al.
Pubblicazione: (2025)
Synthetic training set generation using text-to-audio models for environmental sound classification
di: Ronchini, Francesca, et al.
Pubblicazione: (2024)
di: Ronchini, Francesca, et al.
Pubblicazione: (2024)
Angular Distance Distribution Loss for Audio Classification
di: Almudévar, Antonio, et al.
Pubblicazione: (2024)
di: Almudévar, Antonio, et al.
Pubblicazione: (2024)
A Phoneme-Scale Assessment of Multichannel Speech Enhancement Algorithms
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2024)
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2024)
Evaluating Multichannel Speech Enhancement Algorithms at the Phoneme Scale Across Genders
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2025)
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2025)
Frequency-aware convolution for sound event detection
di: Song, Tao, et al.
Pubblicazione: (2024)
di: Song, Tao, et al.
Pubblicazione: (2024)
PAGURI: a user experience study of creative interaction with text-to-music models
di: Ronchini, Francesca, et al.
Pubblicazione: (2024)
di: Ronchini, Francesca, et al.
Pubblicazione: (2024)
Energy Consumption Trends in Sound Event Detection Systems
di: Douwes, Constance, et al.
Pubblicazione: (2024)
di: Douwes, Constance, et al.
Pubblicazione: (2024)
Domain-Invariant Representation Learning of Bird Sounds
di: Moummad, Ilyass, et al.
Pubblicazione: (2024)
di: Moummad, Ilyass, et al.
Pubblicazione: (2024)
Tracking of Intermittent and Moving Speakers : Dataset and Metrics
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
Fine-tune the pretrained ATST model for sound event detection
di: Shao, Nian, et al.
Pubblicazione: (2023)
di: Shao, Nian, et al.
Pubblicazione: (2023)
Onset and offset weighted loss function for sound event detection
di: Song, Tao
Pubblicazione: (2024)
di: Song, Tao
Pubblicazione: (2024)
Full-frequency dynamic convolution: a physical frequency-dependent convolution for sound event detection
di: Yue, Haobo, et al.
Pubblicazione: (2024)
di: Yue, Haobo, et al.
Pubblicazione: (2024)
Latent Watermarking of Audio Generative Models
di: Roman, Robin San, et al.
Pubblicazione: (2024)
di: Roman, Robin San, et al.
Pubblicazione: (2024)
A decade of DCASE: Achievements, practices, evaluations and future challenges
di: Mesaros, Annamaria, et al.
Pubblicazione: (2024)
di: Mesaros, Annamaria, et al.
Pubblicazione: (2024)
Self-Supervised Learning for Few-Shot Bird Sound Classification
di: Moummad, Ilyass, et al.
Pubblicazione: (2023)
di: Moummad, Ilyass, et al.
Pubblicazione: (2023)
Regularized Contrastive Pre-training for Few-shot Bioacoustic Sound Detection
di: Moummad, Ilyass, et al.
Pubblicazione: (2023)
di: Moummad, Ilyass, et al.
Pubblicazione: (2023)
Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling
di: Gao, Wenmiao, et al.
Pubblicazione: (2025)
di: Gao, Wenmiao, et al.
Pubblicazione: (2025)
Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
di: Vo, Quoc Thinh, et al.
Pubblicazione: (2025)
di: Vo, Quoc Thinh, et al.
Pubblicazione: (2025)
Frequency-Weighted Training Losses for Phoneme-Level DNN-based Speech Enhancement
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2025)
di: Monir, Nasser-Eddine, et al.
Pubblicazione: (2025)
MambaFoley: Foley Sound Generation using Selective State-Space Models
di: Colombo, Marco Furio, et al.
Pubblicazione: (2024)
di: Colombo, Marco Furio, et al.
Pubblicazione: (2024)
Representational learning for an anomalous sound detection system with source separation model
di: Shin, Seunghyeon, et al.
Pubblicazione: (2024)
di: Shin, Seunghyeon, et al.
Pubblicazione: (2024)
Robust detection of overlapping bioacoustic sound events
di: Mahon, Louis, et al.
Pubblicazione: (2025)
di: Mahon, Louis, et al.
Pubblicazione: (2025)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
Synthetic data enables context-aware bioacoustic sound event detection
di: Hoffman, Benjamin, et al.
Pubblicazione: (2025)
di: Hoffman, Benjamin, et al.
Pubblicazione: (2025)
Mixture of Mixups for Multi-label Classification of Rare Anuran Sounds
di: Moummad, Ilyass, et al.
Pubblicazione: (2024)
di: Moummad, Ilyass, et al.
Pubblicazione: (2024)
Posterior Transition Modeling for Unsupervised Diffusion-Based Speech Enhancement
di: Sadeghi, Mostafa, et al.
Pubblicazione: (2025)
di: Sadeghi, Mostafa, et al.
Pubblicazione: (2025)
EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
di: Chu, Yun, et al.
Pubblicazione: (2025)
di: Chu, Yun, et al.
Pubblicazione: (2025)
Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
di: Yasuda, Masahiro, et al.
Pubblicazione: (2025)
di: Yasuda, Masahiro, et al.
Pubblicazione: (2025)
Serial-OE: Anomalous sound detection based on serial method with outlier exposure capable of using small amounts of anomalous data for training
di: Kuroyanagi, Ibuki, et al.
Pubblicazione: (2025)
di: Kuroyanagi, Ibuki, et al.
Pubblicazione: (2025)
AI-Assisted Music Production: A User Study on Text-to-Music Models
di: Ronchini, Francesca, et al.
Pubblicazione: (2025)
di: Ronchini, Francesca, et al.
Pubblicazione: (2025)
Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4
di: Son, Sang Won, et al.
Pubblicazione: (2024)
di: Son, Sang Won, et al.
Pubblicazione: (2024)
Differentiable physics for sound field reconstruction
di: Verburg, Samuel A., et al.
Pubblicazione: (2025)
di: Verburg, Samuel A., et al.
Pubblicazione: (2025)
Some clues to build a sound analysis relevant to hearing
di: Millot, Laurent
Pubblicazione: (2024)
di: Millot, Laurent
Pubblicazione: (2024)
Documenti analoghi
-
A benchmark of state-of-the-art sound event detection systems evaluated on synthetic soundscapes
di: Ronchini, Francesca, et al.
Pubblicazione: (2022) -
The impact of non-target events in synthetic soundscapes for sound event detection
di: Ronchini, Francesca, et al.
Pubblicazione: (2021) -
Description and analysis of novelties introduced in DCASE Task 4 2022 on the baseline system
di: Ronchini, Francesca, et al.
Pubblicazione: (2022) -
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
di: Ronchini, Francesca, et al.
Pubblicazione: (2020) -
Diffused Responsibility: Analyzing the Energy Consumption of Generative Text-to-Audio Diffusion Models
di: Passoni, Riccardo, et al.
Pubblicazione: (2025)