Multitask frame-level learning for few-shot sound event detection
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Zou, Liang, Yan, Genwei, Wang, Ruoyu, Du, Jun, Lei, Meng, Gao, Tian, Fang, Xin |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Frequency-aware convolution for sound event detection
par: Song, Tao, et autres
Publié: (2024)
par: Song, Tao, et autres
Publié: (2024)
Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling
par: Gao, Wenmiao, et autres
Publié: (2025)
par: Gao, Wenmiao, et autres
Publié: (2025)
Onset and offset weighted loss function for sound event detection
par: Song, Tao
Publié: (2024)
par: Song, Tao
Publié: (2024)
Fine-tune the pretrained ATST model for sound event detection
par: Shao, Nian, et autres
Publié: (2023)
par: Shao, Nian, et autres
Publié: (2023)
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
par: Gagnere, Antonin, et autres
Publié: (2024)
par: Gagnere, Antonin, et autres
Publié: (2024)
Full-frequency dynamic convolution: a physical frequency-dependent convolution for sound event detection
par: Yue, Haobo, et autres
Publié: (2024)
par: Yue, Haobo, et autres
Publié: (2024)
Performance and energy balance: a comprehensive study of state-of-the-art sound event detection systems
par: Ronchini, Francesca, et autres
Publié: (2023)
par: Ronchini, Francesca, et autres
Publié: (2023)
Representational learning for an anomalous sound detection system with source separation model
par: Shin, Seunghyeon, et autres
Publié: (2024)
par: Shin, Seunghyeon, et autres
Publié: (2024)
Robust detection of overlapping bioacoustic sound events
par: Mahon, Louis, et autres
Publié: (2025)
par: Mahon, Louis, et autres
Publié: (2025)
Resnet-conformer network with shared weights and attention mechanism for sound event localization, detection, and distance estimation
par: Vo, Quoc Thinh, et autres
Publié: (2025)
par: Vo, Quoc Thinh, et autres
Publié: (2025)
The impact of non-target events in synthetic soundscapes for sound event detection
par: Ronchini, Francesca, et autres
Publié: (2021)
par: Ronchini, Francesca, et autres
Publié: (2021)
Incorporating Spatial Cues in Modular Speaker Diarization for Multi-channel Multi-party Meetings
par: Wang, Ruoyu, et autres
Publié: (2024)
par: Wang, Ruoyu, et autres
Publié: (2024)
Multi-level Temporal-channel Speaker Retrieval for Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2023)
par: Wang, Zhichao, et autres
Publié: (2023)
Synthetic data enables context-aware bioacoustic sound event detection
par: Hoffman, Benjamin, et autres
Publié: (2025)
par: Hoffman, Benjamin, et autres
Publié: (2025)
Towards Building Speech Large Language Models for Multitask Understanding in Low-Resource Languages
par: Shao, Mingchen, et autres
Publié: (2025)
par: Shao, Mingchen, et autres
Publié: (2025)
A benchmark of state-of-the-art sound event detection systems evaluated on synthetic soundscapes
par: Ronchini, Francesca, et autres
Publié: (2022)
par: Ronchini, Francesca, et autres
Publié: (2022)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
par: Gong, Rong, et autres
Publié: (2024)
par: Gong, Rong, et autres
Publié: (2024)
InsectSet459: an open dataset of insect sounds for bioacoustic machine learning
par: Faiß, Marius, et autres
Publié: (2025)
par: Faiß, Marius, et autres
Publié: (2025)
Improving Noise Robustness of LLM-based Zero-shot TTS via Discrete Acoustic Token Denoising
par: Lu, Ye-Xin, et autres
Publié: (2025)
par: Lu, Ye-Xin, et autres
Publié: (2025)
EZhouNet:A framework based on graph neural network and anchor interval for the respiratory sound event detection
par: Chu, Yun, et autres
Publié: (2025)
par: Chu, Yun, et autres
Publié: (2025)
Efficient learning-based sound propagation for virtual and real-world audio processing applications
par: Ratnarajah, Anton Jeran
Publié: (2024)
par: Ratnarajah, Anton Jeran
Publié: (2024)
CoDiff-VC: A Codec-Assisted Diffusion Model for Zero-shot Voice Conversion
par: Li, Yuke, et autres
Publié: (2024)
par: Li, Yuke, et autres
Publié: (2024)
ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
par: Zhu, Xinfa, et autres
Publié: (2025)
par: Zhu, Xinfa, et autres
Publié: (2025)
Adaptive high-precision sound source localization at low frequencies based on convolutional neural network
par: Ma, Wenbo, et autres
Publié: (2024)
par: Ma, Wenbo, et autres
Publié: (2024)
Serial-OE: Anomalous sound detection based on serial method with outlier exposure capable of using small amounts of anomalous data for training
par: Kuroyanagi, Ibuki, et autres
Publié: (2025)
par: Kuroyanagi, Ibuki, et autres
Publié: (2025)
StreamVoice+: Evolving into End-to-end Streaming Zero-shot Voice Conversion
par: Wang, Zhichao, et autres
Publié: (2024)
par: Wang, Zhichao, et autres
Publié: (2024)
Sound event localization and detection based on crnn using rectangular filters and channel rotation data augmentation
par: Ronchini, Francesca, et autres
Publié: (2020)
par: Ronchini, Francesca, et autres
Publié: (2020)
DCF-DS: Deep Cascade Fusion of Diarization and Separation for Speech Recognition under Realistic Single-Channel Conditions
par: Niu, Shu-Tong, et autres
Publié: (2024)
par: Niu, Shu-Tong, et autres
Publié: (2024)
Sound event detection based on auxiliary decoder and maximum probability aggregation for DCASE Challenge 2024 Task 4
par: Son, Sang Won, et autres
Publié: (2024)
par: Son, Sang Won, et autres
Publié: (2024)
Fine-grained Preference Optimization Improves Zero-shot Text-to-Speech
par: Yao, Jixun, et autres
Publié: (2025)
par: Yao, Jixun, et autres
Publié: (2025)
Learning to detect an animal sound from five examples
par: Nolasco, Inês, et autres
Publié: (2023)
par: Nolasco, Inês, et autres
Publié: (2023)
Differentiable physics for sound field reconstruction
par: Verburg, Samuel A., et autres
Publié: (2025)
par: Verburg, Samuel A., et autres
Publié: (2025)
Improving Audio-Text Retrieval via Hierarchical Cross-Modal Interaction and Auxiliary Captions
par: Xin, Yifei, et autres
Publié: (2023)
par: Xin, Yifei, et autres
Publié: (2023)
Experimental Results of Underwater Sound Speed Profile Inversion by Few-shot Multi-task Learning
par: Huang, Wei, et autres
Publié: (2023)
par: Huang, Wei, et autres
Publié: (2023)
XWSB: A Blend System Utilizing XLS-R and WavLM with SLS Classifier detection system for SVDD 2024 Challenge
par: Zhang, Qishan, et autres
Publié: (2024)
par: Zhang, Qishan, et autres
Publié: (2024)
SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis
par: Wang, Helin, et autres
Publié: (2024)
par: Wang, Helin, et autres
Publié: (2024)
An Enhanced Audio Feature Tailored for Anomalous Sound Detection Based on Pre-trained Models
par: Zhong, Guirui, et autres
Publié: (2025)
par: Zhong, Guirui, et autres
Publié: (2025)
Enhancing Zero-shot Audio Classification using Sound Attribute Knowledge from Large Language Models
par: Xu, Xuenan, et autres
Publié: (2024)
par: Xu, Xuenan, et autres
Publié: (2024)
The Neural-SRP method for positional sound source localization
par: Grinstein, Eric, et autres
Publié: (2024)
par: Grinstein, Eric, et autres
Publié: (2024)
The USTC-NERCSLIP Systems for the CHiME-8 NOTSOFAR-1 Challenge
par: Niu, Shutong, et autres
Publié: (2024)
par: Niu, Shutong, et autres
Publié: (2024)
Documents similaires
-
Frequency-aware convolution for sound event detection
par: Song, Tao, et autres
Publié: (2024) -
Stereo sound event localization and detection based on PSELDnet pretraining and BiMamba sequence modeling
par: Gao, Wenmiao, et autres
Publié: (2025) -
Onset and offset weighted loss function for sound event detection
par: Song, Tao
Publié: (2024) -
Fine-tune the pretrained ATST model for sound event detection
par: Shao, Nian, et autres
Publié: (2023) -
A Contrastive Self-Supervised Learning scheme for beat tracking amenable to few-shot learning
par: Gagnere, Antonin, et autres
Publié: (2024)