Formula-Supervised Sound Event Detection: Pre-Training Without Real Data
Fuente:
arXiv
Guardado en:
| Autores principales: | Shibata, Yuto, Tanaka, Keitaro, Bando, Yoshiaki, Imoto, Keisuke, Kataoka, Hirokatsu, Aoki, Yoshimitsu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
por: Fujita, Yoto, et al.
Publicado: (2024)
por: Fujita, Yoto, et al.
Publicado: (2024)
LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
por: Koga, Naoki, et al.
Publicado: (2024)
por: Koga, Naoki, et al.
Publicado: (2024)
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
por: Imoto, Keisuke
Publicado: (2025)
por: Imoto, Keisuke
Publicado: (2025)
BGM2Pose: Active 3D Human Pose Estimation with Non-Stationary Sounds
por: Shibata, Yuto, et al.
Publicado: (2025)
por: Shibata, Yuto, et al.
Publicado: (2025)
Acoustic-based 3D Human Pose Estimation Robust to Human Position
por: Oumi, Yusuke, et al.
Publicado: (2024)
por: Oumi, Yusuke, et al.
Publicado: (2024)
Sign-to-Speech Prosody Transfer via Sign Reconstruction-based GAN
por: Manabe, Toranosuke, et al.
Publicado: (2026)
por: Manabe, Toranosuke, et al.
Publicado: (2026)
Towards Open World Sound Event Detection
por: Hai, P. H., et al.
Publicado: (2026)
por: Hai, P. H., et al.
Publicado: (2026)
SHAMaNS: Sound Localization with Hybrid Alpha-Stable Spatial Measure and Neural Steerer
por: Di Carlo, Diego, et al.
Publicado: (2025)
por: Di Carlo, Diego, et al.
Publicado: (2025)
Sound Scene Synthesis at the DCASE 2024 Challenge
por: Lagrange, Mathieu, et al.
Publicado: (2025)
por: Lagrange, Mathieu, et al.
Publicado: (2025)
Prototype based Masked Audio Model for Self-Supervised Learning of Sound Event Detection
por: Cai, Pengfei, et al.
Publicado: (2024)
por: Cai, Pengfei, et al.
Publicado: (2024)
Enhanced Sound Event Localization and Detection in Real 360-degree audio-visual soundscapes
por: Roman, Adrian S., et al.
Publicado: (2024)
por: Roman, Adrian S., et al.
Publicado: (2024)
Improving Anomalous Sound Detection with Attribute-aware Representation from Domain-adaptive Pre-training
por: Fang, Xin, et al.
Publicado: (2025)
por: Fang, Xin, et al.
Publicado: (2025)
MoireDB: Formula-generated Interference-fringe Image Dataset
por: Matsuo, Yuto, et al.
Publicado: (2025)
por: Matsuo, Yuto, et al.
Publicado: (2025)
MAT-SED: A Masked Audio Transformer with Masked-Reconstruction Based Pre-training for Sound Event Detection
por: Cai, Pengfei, et al.
Publicado: (2024)
por: Cai, Pengfei, et al.
Publicado: (2024)
Detect Any Sound: Open-Vocabulary Sound Event Detection with Multi-Modal Queries
por: Cai, Pengfei, et al.
Publicado: (2025)
por: Cai, Pengfei, et al.
Publicado: (2025)
Leveraging Language Model Capabilities for Sound Event Detection
por: Wang, Hualei, et al.
Publicado: (2023)
por: Wang, Hualei, et al.
Publicado: (2023)
Improving Anomalous Sound Detection via Low-Rank Adaptation Fine-Tuning of Pre-Trained Audio Models
por: Zheng, Xinhu, et al.
Publicado: (2024)
por: Zheng, Xinhu, et al.
Publicado: (2024)
The Sounds of Home: A Speech-Removed Residential Audio Dataset for Sound Event Detection
por: Bibbó, Gabriel, et al.
Publicado: (2024)
por: Bibbó, Gabriel, et al.
Publicado: (2024)
Pianist Transformer: Towards Expressive Piano Performance Rendering via Scalable Self-Supervised Pre-Training
por: You, Hong-Jie, et al.
Publicado: (2025)
por: You, Hong-Jie, et al.
Publicado: (2025)
Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods
por: Zhou, Xuanru, et al.
Publicado: (2026)
por: Zhou, Xuanru, et al.
Publicado: (2026)
FlexSED: Towards Open-Vocabulary Sound Event Detection
por: Hai, Jiarui, et al.
Publicado: (2025)
por: Hai, Jiarui, et al.
Publicado: (2025)
Studying the Effect of Audio Filters in Pre-Trained Models for Environmental Sound Classification
por: Dawn, Aditya, et al.
Publicado: (2024)
por: Dawn, Aditya, et al.
Publicado: (2024)
Challenge on Sound Scene Synthesis: Evaluating Text-to-Audio Generation
por: Lee, Junwon, et al.
Publicado: (2024)
por: Lee, Junwon, et al.
Publicado: (2024)
Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling
por: Li, Xingyuan, et al.
Publicado: (2026)
por: Li, Xingyuan, et al.
Publicado: (2026)
'Studies for': A Human-AI Co-Creative Sound Artwork Using a Real-time Multi-channel Sound Generation Model
por: Nagashima, Chihiro, et al.
Publicado: (2025)
por: Nagashima, Chihiro, et al.
Publicado: (2025)
Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
por: Truong, Duc-Tuan, et al.
Publicado: (2025)
por: Truong, Duc-Tuan, et al.
Publicado: (2025)
Discrete Speech Unit Extraction via Independent Component Analysis
por: Nakamura, Tomohiko, et al.
Publicado: (2025)
por: Nakamura, Tomohiko, et al.
Publicado: (2025)
Environmental Sound Deepfake Detection Using Deep-Learning Framework
por: Pham, Lam, et al.
Publicado: (2026)
por: Pham, Lam, et al.
Publicado: (2026)
Rethinking Leveraging Pre-Trained Multi-Layer Representations for Speaker Verification
por: Kim, Jin Sob, et al.
Publicado: (2025)
por: Kim, Jin Sob, et al.
Publicado: (2025)
MoireMix: A Formula-Based Data Augmentation for Improving Image Classification Robustness
por: Matsuo, Yuto, et al.
Publicado: (2026)
por: Matsuo, Yuto, et al.
Publicado: (2026)
How Much Does Machine Identity Matter in Anomalous Sound Detection at Test Time?
por: Wilkinghoff, Kevin, et al.
Publicado: (2026)
por: Wilkinghoff, Kevin, et al.
Publicado: (2026)
Sub-Band Spectral Matching with Localized Score Aggregation for Robust Anomalous Sound Detection
por: Saengthong, Phurich, et al.
Publicado: (2026)
por: Saengthong, Phurich, et al.
Publicado: (2026)
SELD-Mamba: Selective State-Space Model for Sound Event Localization and Detection with Source Distance Estimation
por: Mu, Da, et al.
Publicado: (2024)
por: Mu, Da, et al.
Publicado: (2024)
Run-Time Adaptation of Neural Beamforming for Robust Speech Dereverberation and Denoising
por: Fujita, Yoto, et al.
Publicado: (2024)
por: Fujita, Yoto, et al.
Publicado: (2024)
Pediatric Asthma Detection with Googles HeAR Model: An AI-Driven Respiratory Sound Classifier
por: Ehtesham, Abul, et al.
Publicado: (2025)
por: Ehtesham, Abul, et al.
Publicado: (2025)
TopSeg: A Multi-Scale Topological Framework for Data-Efficient Heart Sound Segmentation
por: Zhang, Peihong, et al.
Publicado: (2025)
por: Zhang, Peihong, et al.
Publicado: (2025)
Hybrid Disagreement-Diversity Active Learning for Bioacoustic Sound Event Detection
por: Zhang, Shiqi, et al.
Publicado: (2025)
por: Zhang, Shiqi, et al.
Publicado: (2025)
Universal Sound Separation with Self-Supervised Audio Masked Autoencoder
por: Zhao, Junqi, et al.
Publicado: (2024)
por: Zhao, Junqi, et al.
Publicado: (2024)
Machine Anomalous Sound Detection Using Spectral-temporal Modulation Representations Derived from Machine-specific Filterbanks
por: Li, Kai, et al.
Publicado: (2024)
por: Li, Kai, et al.
Publicado: (2024)
SynSonic: Augmenting Sound Event Detection through Text-to-Audio Diffusion ControlNet and Effective Sample Filtering
por: Hai, Jiarui, et al.
Publicado: (2025)
por: Hai, Jiarui, et al.
Publicado: (2025)
Ejemplares similares
-
DOA-Aware Audio-Visual Self-Supervised Learning for Sound Event Localization and Detection
por: Fujita, Yoto, et al.
Publicado: (2024) -
LEAD Dataset: How Can Labels for Sound Event Detection Vary Depending on Annotators?
por: Koga, Naoki, et al.
Publicado: (2024) -
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
por: Imoto, Keisuke
Publicado: (2025) -
BGM2Pose: Active 3D Human Pose Estimation with Non-Stationary Sounds
por: Shibata, Yuto, et al.
Publicado: (2025) -
Acoustic-based 3D Human Pose Estimation Robust to Human Position
por: Oumi, Yusuke, et al.
Publicado: (2024)