CST-former: Multidimensional Attention-based Transformer for Sound Event Localization and Detection in Real Scenes
Fuente:
arXiv
Saved in:
| Main Authors: | Shul, Yusun, Choi, Dayun, Choi, Jung-Woo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
by: Choi, Dayun, et al.
Published: (2025)
by: Choi, Dayun, et al.
Published: (2025)
Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues
by: Choi, Dayun, et al.
Published: (2024)
by: Choi, Dayun, et al.
Published: (2024)
Sound Separation and Classification with Object and Semantic Guidance
by: Kwon, Younghoo, et al.
Published: (2025)
by: Kwon, Younghoo, et al.
Published: (2025)
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
by: Lee, Dongheon, et al.
Published: (2024)
by: Lee, Dongheon, et al.
Published: (2024)
Self-Guided Target Sound Extraction and Classification Through Universal Sound Separation Model and Multiple Clues
by: Kwon, Younghoo, et al.
Published: (2025)
by: Kwon, Younghoo, et al.
Published: (2025)
SEMamba++: A General Speech Restoration Framework Leveraging Global, Local, and Periodic Spectral Patterns
by: Lee, Yongjoon, et al.
Published: (2026)
by: Lee, Yongjoon, et al.
Published: (2026)
DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis
by: Lee, Dongheon, et al.
Published: (2025)
by: Lee, Dongheon, et al.
Published: (2025)
Speech-Declipping Transformer with Complex Spectrogram and Learnerble Temporal Features
by: Kwon, Younghoo, et al.
Published: (2024)
by: Kwon, Younghoo, et al.
Published: (2024)
Neural acoustic multipole splatting for room impulse response synthesis
by: Baek, Geonwoo, et al.
Published: (2025)
by: Baek, Geonwoo, et al.
Published: (2025)
MVANet: Multi-Stage Video Attention Network for Sound Event Localization and Detection with Source Distance Estimation
by: Hong, Hengyi, et al.
Published: (2024)
by: Hong, Hengyi, et al.
Published: (2024)
FlowW2N: Whispered-to-Normal Speech Conversion via Flow-Matching
by: Ritter-Gutierrez, Fabian, et al.
Published: (2026)
by: Ritter-Gutierrez, Fabian, et al.
Published: (2026)
SwG-former: A Sliding-Window Graph Convolutional Network for Simultaneous Spatial-Temporal Information Extraction in Sound Event Localization and Detection
by: Huang, Weiming, et al.
Published: (2023)
by: Huang, Weiming, et al.
Published: (2023)
Class-Incremental Learning for Sound Event Localization and Detection
by: Pandey, Ruchi, et al.
Published: (2024)
by: Pandey, Ruchi, et al.
Published: (2024)
DeFTAN-II: Efficient Multichannel Speech Enhancement with Subgroup Processing
by: Lee, Dongheon, et al.
Published: (2023)
by: Lee, Dongheon, et al.
Published: (2023)
DISPATCH: Distilling Selective Patches for Speech Enhancement
by: Kim, Dohwan, et al.
Published: (2025)
by: Kim, Dohwan, et al.
Published: (2025)
3D Room Geometry Inference from Multichannel Room Impulse Response using Deep Neural Network
by: Yeon, Inmo, et al.
Published: (2024)
by: Yeon, Inmo, et al.
Published: (2024)
RAF: Relativistic Adversarial Feedback For Universal Speech Synthesis
by: Lee, Yongjoon, et al.
Published: (2026)
by: Lee, Yongjoon, et al.
Published: (2026)
Binaural Sound Event Localization and Detection based on HRTF Cues for Humanoid Robots
by: Lee, Gyeong-Tae, et al.
Published: (2025)
by: Lee, Gyeong-Tae, et al.
Published: (2025)
Zero- and Few-shot Sound Event Localization and Detection
by: Shimada, Kazuki, et al.
Published: (2023)
by: Shimada, Kazuki, et al.
Published: (2023)
Self Training and Ensembling Frequency Dependent Networks with Coarse Prediction Pooling and Sound Event Bounding Boxes
by: Nam, Hyeonuk, et al.
Published: (2024)
by: Nam, Hyeonuk, et al.
Published: (2024)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
by: Berghi, Davide, et al.
Published: (2025)
by: Berghi, Davide, et al.
Published: (2025)
Text-Queried Target Sound Event Localization
by: Zhao, Jinzheng, et al.
Published: (2024)
by: Zhao, Jinzheng, et al.
Published: (2024)
Binaural Sound Event Localization and Detection Neural Network based on HRTF Localization Cues for Humanoid Robots
by: Lee, Gyeong-Tae
Published: (2025)
by: Lee, Gyeong-Tae
Published: (2025)
Temporal Attention Pooling for Frequency Dynamic Convolution in Sound Event Detection
by: Nam, Hyeonuk, et al.
Published: (2025)
by: Nam, Hyeonuk, et al.
Published: (2025)
Location-Oriented Sound Event Localization and Detection with Spatial Mapping and Regression Localization
by: Zhang, Xueping, et al.
Published: (2025)
by: Zhang, Xueping, et al.
Published: (2025)
Effective Pre-Training of Audio Transformers for Sound Event Detection
by: Schmid, Florian, et al.
Published: (2024)
by: Schmid, Florian, et al.
Published: (2024)
Speech Emotion Recognition Via CNN-Transformer and Multidimensional Attention Mechanism
by: Tang, Xiaoyu, et al.
Published: (2024)
by: Tang, Xiaoyu, et al.
Published: (2024)
Inter-channel Conv-TasNet for multichannel speech enhancement
by: Lee, Dongheon, et al.
Published: (2021)
by: Lee, Dongheon, et al.
Published: (2021)
Multi-Iteration Multi-Stage Fine-Tuning of Transformers for Sound Event Detection with Heterogeneous Datasets
by: Schmid, Florian, et al.
Published: (2024)
by: Schmid, Florian, et al.
Published: (2024)
String Sound Synthesizer on GPU-accelerated Finite Difference Scheme
by: Lee, Jin Woo, et al.
Published: (2023)
by: Lee, Jin Woo, et al.
Published: (2023)
Joint Analysis of Acoustic Scenes and Sound Events Based on Semi-Supervised Training of Sound Events With Partial Labels
by: Imoto, Keisuke
Published: (2025)
by: Imoto, Keisuke
Published: (2025)
JiTTER: Jigsaw Temporal Transformer for Event Reconstruction for Self-Supervised Sound Event Detection
by: Nam, Hyeonuk, et al.
Published: (2025)
by: Nam, Hyeonuk, et al.
Published: (2025)
BTS: Bridging Text and Sound Modalities for Metadata-Aided Respiratory Sound Classification
by: Kim, June-Woo, et al.
Published: (2024)
by: Kim, June-Woo, et al.
Published: (2024)
Learning How to Listen: A Temporal-Frequential Attention Model for Sound Event Detection
by: Shen, Yu-Han, et al.
Published: (2018)
by: Shen, Yu-Han, et al.
Published: (2018)
Generating Diverse Audio-Visual 360 Soundscapes for Sound Event Localization and Detection
by: Roman, Adrian S., et al.
Published: (2025)
by: Roman, Adrian S., et al.
Published: (2025)
A Two-Step Learning Framework for Enhancing Sound Event Localization and Detection
by: Yu, Hogeon
Published: (2025)
by: Yu, Hogeon
Published: (2025)
Selective-Memory Meta-Learning with Environment Representations for Sound Event Localization and Detection
by: Hu, Jinbo, et al.
Published: (2023)
by: Hu, Jinbo, et al.
Published: (2023)
Enhancing Situational Awareness in Wearable Audio Devices Using a Lightweight Sound Event Localization and Detection System
by: Yeow, Jun-Wei, et al.
Published: (2025)
by: Yeow, Jun-Wei, et al.
Published: (2025)
Squeeze-and-Excite ResNet-Conformers for Sound Event Localization, Detection, and Distance Estimation for DCASE 2024 Challenge
by: Yeow, Jun Wei, et al.
Published: (2024)
by: Yeow, Jun Wei, et al.
Published: (2024)
RGI-Net: 3D Room Geometry Inference from Room Impulse Responses With Hidden First-Order Reflections
by: Yeon, Inmo, et al.
Published: (2023)
by: Yeon, Inmo, et al.
Published: (2023)
Similar Items
-
SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
by: Choi, Dayun, et al.
Published: (2025) -
Multichannel-to-Multichannel Target Sound Extraction Using Direction and Timestamp Clues
by: Choi, Dayun, et al.
Published: (2024) -
Sound Separation and Classification with Object and Semantic Guidance
by: Kwon, Younghoo, et al.
Published: (2025) -
DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification
by: Lee, Dongheon, et al.
Published: (2024) -
Self-Guided Target Sound Extraction and Classification Through Universal Sound Separation Model and Multiple Clues
by: Kwon, Younghoo, et al.
Published: (2025)