Sound Event Detection with Boundary-Aware Optimization and Inference
Fuente:
arXiv
Saved in:
| Main Authors: | Schmid, Florian, Tang, Chi Ian, Parekh, Sanjeel, Ithapu, Vamsi Krishna, Ortiz, Juan Azcarreta, Ferroni, Giacomo, Qian, Yijun, Jasonas, Arnoldas, Frateanu, Cosmin, Clark, Camilla, Widmer, Gerhard, Bilen, Çağdaş |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
More Than A Shortcut: A Hyperbolic Approach To Early-Exit Networks
by: Bhosale, Swapnil, et al.
Published: (2025)
by: Bhosale, Swapnil, et al.
Published: (2025)
Efficient Neural and Numerical Methods for High-Quality Online Speech Spectrogram Inversion via Gradient Theorem
by: Fernandez, Andres, et al.
Published: (2025)
by: Fernandez, Andres, et al.
Published: (2025)
Text-to-Stage: Spatial Layouts from Long-form Narratives
by: Hernandez, Jefferson, et al.
Published: (2026)
by: Hernandez, Jefferson, et al.
Published: (2026)
ArrayDPS-Refine: Generative Refinement of Discriminative Multi-Channel Speech Enhancement
by: Xu, Zhongweiyang, et al.
Published: (2026)
by: Xu, Zhongweiyang, et al.
Published: (2026)
Learning to Highlight Audio by Watching Movies
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Spatial-Magnifier: Spatial upsampling for multichannel speech enhancement
by: Lee, Dongheon, et al.
Published: (2026)
by: Lee, Dongheon, et al.
Published: (2026)
Unified Diffusion Refinement for Multi-Channel Speech Enhancement and Separation
by: Xu, Zhongweiyang, et al.
Published: (2026)
by: Xu, Zhongweiyang, et al.
Published: (2026)
Exploring Performance-Complexity Trade-Offs in Sound Event Detection Models
by: Morocutti, Tobias, et al.
Published: (2025)
by: Morocutti, Tobias, et al.
Published: (2025)
Improving Audio Spectrogram Transformers for Sound Event Detection Through Multi-Stage Training
by: Schmid, Florian, et al.
Published: (2024)
by: Schmid, Florian, et al.
Published: (2024)
Effective Pre-Training of Audio Transformers for Sound Event Detection
by: Schmid, Florian, et al.
Published: (2024)
by: Schmid, Florian, et al.
Published: (2024)
Multi-Iteration Multi-Stage Fine-Tuning of Transformers for Sound Event Detection with Heterogeneous Datasets
by: Schmid, Florian, et al.
Published: (2024)
by: Schmid, Florian, et al.
Published: (2024)
Controlling the Parameterized Multi-channel Wiener Filter using a tiny neural network
by: Grinstein, Eric, et al.
Published: (2025)
by: Grinstein, Eric, et al.
Published: (2025)
Surgery for lung cancer as the second primary malignancy
by: Arnoldas Krasauskas
Published: (2012)
by: Arnoldas Krasauskas
Published: (2012)
Modulating State Space Model with SlowFast Framework for Compute-Efficient Ultra Low-Latency Speech Enhancement
by: Cheng, Longbiao, et al.
Published: (2024)
by: Cheng, Longbiao, et al.
Published: (2024)
Hearing Loss Detection from Facial Expressions in One-on-one Conversations
by: Yin, Yufeng, et al.
Published: (2024)
by: Yin, Yufeng, et al.
Published: (2024)
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception
by: Chowdhury, Sanjoy, et al.
Published: (2025)
by: Chowdhury, Sanjoy, et al.
Published: (2025)
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
by: Jia, Wenqi, et al.
Published: (2023)
by: Jia, Wenqi, et al.
Published: (2023)
TACOS: Temporally-aligned Audio CaptiOnS for Language-Audio Pretraining
by: Primus, Paul, et al.
Published: (2025)
by: Primus, Paul, et al.
Published: (2025)
Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval
by: Primus, Paul, et al.
Published: (2024)
by: Primus, Paul, et al.
Published: (2024)
Network Engineering in the Era of AI: A Technical Review
by: Vamsi Krishna Gadireddy
Published: (2025)
by: Vamsi Krishna Gadireddy
Published: (2025)
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos
by: Yun, Heeseung, et al.
Published: (2024)
by: Yun, Heeseung, et al.
Published: (2024)
Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment
by: Hong, Joanna, et al.
Published: (2025)
by: Hong, Joanna, et al.
Published: (2025)
Executable Boundary Contracts for Sound Event Traces
by: Alpay, Faruk, et al.
Published: (2026)
by: Alpay, Faruk, et al.
Published: (2026)
Effects of different protein sources on growth performance and food consumption of goldfish, Carassius auratus
by: Bilen, S., et al.
Published: (2013)
by: Bilen, S., et al.
Published: (2013)
Improving Query-by-Vocal Imitation with Contrastive Learning and Audio Pretraining
by: Greif, Jonathan, et al.
Published: (2024)
by: Greif, Jonathan, et al.
Published: (2024)
Creating a Good Teacher for Knowledge Distillation in Acoustic Scene Classification
by: Morocutti, Tobias, et al.
Published: (2025)
by: Morocutti, Tobias, et al.
Published: (2025)
Device-Robust Acoustic Scene Classification via Impulse Response Augmentation
by: Morocutti, Tobias, et al.
Published: (2023)
by: Morocutti, Tobias, et al.
Published: (2023)
Trivance: Latency-Optimal AllReduce by Shortcutting Multiport Networks
by: Juerss, Anton, et al.
Published: (2026)
by: Juerss, Anton, et al.
Published: (2026)
Understanding the Throughput Bounds of Reconfigurable Datacenter Networks
by: Addanki, Vamsi, et al.
Published: (2024)
by: Addanki, Vamsi, et al.
Published: (2024)
Credence: Augmenting Datacenter Switch Buffer Sharing with ML Predictions
by: Addanki, Vamsi, et al.
Published: (2024)
by: Addanki, Vamsi, et al.
Published: (2024)
Conditional Flow Matching for Visually-Guided Acoustic Highlighting
by: Malard, Hugo, et al.
Published: (2026)
by: Malard, Hugo, et al.
Published: (2026)
Hearing Anywhere in Any Environment
by: Liu, Xiulong, et al.
Published: (2025)
by: Liu, Xiulong, et al.
Published: (2025)
A Benchmark Time Series Dataset for Semiconductor Fabrication Manufacturing Constructed using Component-based Discrete-Event Simulation Models
by: Pendyala, Vamsi Krishna, et al.
Published: (2024)
by: Pendyala, Vamsi Krishna, et al.
Published: (2024)
Rule-based outlier detection of AI-generated anatomy segmentations
by: Krishnaswamy, Deepa, et al.
Published: (2024)
by: Krishnaswamy, Deepa, et al.
Published: (2024)
The effects of oxygen supplementation on growth and survival of rainbow trout (Oncorhynchus mykiss) in different stocking densities
by: Bilen, S., et al.
Published: (2015)
by: Bilen, S., et al.
Published: (2015)
Mental Events
by: Schmid, Wolf
Published: (2022)
by: Schmid, Wolf
Published: (2022)
Physics-Informed Neural Network Approaches for Sparse Data Flow Reconstruction of Unsteady Flow Around Complex Geometries
by: Malineni, Vamsi Sai Krishna, et al.
Published: (2025)
by: Malineni, Vamsi Sai Krishna, et al.
Published: (2025)
On Temporal Guidance and Iterative Refinement in Audio Source Separation
by: Morocutti, Tobias, et al.
Published: (2025)
by: Morocutti, Tobias, et al.
Published: (2025)
Ethereal: Divide and Conquer Network Load Balancing in Large-Scale Distributed Training
by: Addanki, Vamsi, et al.
Published: (2024)
by: Addanki, Vamsi, et al.
Published: (2024)
OPTIMIZATION OF CYLINDRICALLY CURVED COMPOSITE PANELS UNDER COMPRESSIVE LOADS
by: S Adali, et al.
Published: (2009)
by: S Adali, et al.
Published: (2009)
Similar Items
-
More Than A Shortcut: A Hyperbolic Approach To Early-Exit Networks
by: Bhosale, Swapnil, et al.
Published: (2025) -
Efficient Neural and Numerical Methods for High-Quality Online Speech Spectrogram Inversion via Gradient Theorem
by: Fernandez, Andres, et al.
Published: (2025) -
Text-to-Stage: Spatial Layouts from Long-form Narratives
by: Hernandez, Jefferson, et al.
Published: (2026) -
ArrayDPS-Refine: Generative Refinement of Discriminative Multi-Channel Speech Enhancement
by: Xu, Zhongweiyang, et al.
Published: (2026) -
Learning to Highlight Audio by Watching Movies
by: Huang, Chao, et al.
Published: (2025)