MHANet: Multi-scale Hybrid Attention Network for Auditory Attention Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Lu, Fan, Cunhang, Zhang, Hongyu, Zhang, Jingjing, Yang, Xiaoke, Zhou, Jian, Lv, Zhao |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
di: Yan, Sheng, et al.
Pubblicazione: (2024)
di: Yan, Sheng, et al.
Pubblicazione: (2024)
Evaluating Spatialized Auditory Cues for Rapid Attention Capture in XR
di: Kim, Yoonsang, et al.
Pubblicazione: (2026)
di: Kim, Yoonsang, et al.
Pubblicazione: (2026)
M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
Auditory Attention Decoding without Spatial Information: A Diotic EEG Study
di: Yoshino, Masahiro, et al.
Pubblicazione: (2026)
di: Yoshino, Masahiro, et al.
Pubblicazione: (2026)
WSCoach: Wearable Real-time Auditory Feedback for Reducing Unwanted Words in Daily Communication
di: Youpeng, Zhang, et al.
Pubblicazione: (2025)
di: Youpeng, Zhang, et al.
Pubblicazione: (2025)
Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
di: Jiang, Xilin, et al.
Pubblicazione: (2025)
di: Jiang, Xilin, et al.
Pubblicazione: (2025)
Spatial Reconstructed Local Attention Res2Net with F0 Subband for Fake Speech Detection
di: Fan, Cunhang, et al.
Pubblicazione: (2023)
di: Fan, Cunhang, et al.
Pubblicazione: (2023)
Single-word Auditory Attention Decoding Using Deep Learning Model
di: Nguyen, Nhan Duc Thanh, et al.
Pubblicazione: (2024)
di: Nguyen, Nhan Duc Thanh, et al.
Pubblicazione: (2024)
DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
Early Detection of Furniture-Infesting Wood-Boring Beetles Using CNN-LSTM Networks and MFCC-Based Acoustic Features
di: Manukalpa, J. M. Chan Sri, et al.
Pubblicazione: (2025)
di: Manukalpa, J. M. Chan Sri, et al.
Pubblicazione: (2025)
SSM2Mel: State Space Model to Reconstruct Mel Spectrogram from the EEG
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
di: Fan, Cunhang, et al.
Pubblicazione: (2025)
AVE Speech: A Comprehensive Multi-Modal Dataset for Speech Recognition Integrating Audio, Visual, and Electromyographic Signals
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
di: Zhou, Dongliang, et al.
Pubblicazione: (2025)
Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors
di: Nowrin, Sadia, et al.
Pubblicazione: (2024)
di: Nowrin, Sadia, et al.
Pubblicazione: (2024)
Interactive Sonification for Health and Energy using ChucK and Unity
di: Zhao, Yichun, et al.
Pubblicazione: (2024)
di: Zhao, Yichun, et al.
Pubblicazione: (2024)
Springboard, Roadblock or "Crutch"?: How Transgender Users Leverage Voice Changers for Gender Presentation in Social Virtual Reality
di: Povinelli, Kassie, et al.
Pubblicazione: (2024)
di: Povinelli, Kassie, et al.
Pubblicazione: (2024)
Beyond-Voice: Towards Continuous 3D Hand Pose Tracking on Commercial Home Assistant Devices
di: Li, Yin, et al.
Pubblicazione: (2023)
di: Li, Yin, et al.
Pubblicazione: (2023)
Subject Disentanglement Neural Network for Speech Envelope Reconstruction from EEG
di: Zhang, Li, et al.
Pubblicazione: (2025)
di: Zhang, Li, et al.
Pubblicazione: (2025)
UltrasonicSpheres: Localized, Multi-Channel Sound Spheres Using Off-the-Shelf Speakers and Earables
di: Küttner, Michael, et al.
Pubblicazione: (2025)
di: Küttner, Michael, et al.
Pubblicazione: (2025)
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
di: Zheng, Naijun, et al.
Pubblicazione: (2025)
di: Zheng, Naijun, et al.
Pubblicazione: (2025)
SingVisio: Visual Analytics of Diffusion Model for Singing Voice Conversion
di: Xue, Liumeng, et al.
Pubblicazione: (2024)
di: Xue, Liumeng, et al.
Pubblicazione: (2024)
Advancing User-Voice Interaction: Exploring Emotion-Aware Voice Assistants Through a Role-Swapping Approach
di: Ma, Yong, et al.
Pubblicazione: (2025)
di: Ma, Yong, et al.
Pubblicazione: (2025)
Improving Multimodal Emotion Recognition by Leveraging Acoustic Adaptation and Visual Alignment
di: Zhao, Zhixian, et al.
Pubblicazione: (2024)
di: Zhao, Zhixian, et al.
Pubblicazione: (2024)
USpeech: Ultrasound-Enhanced Speech with Minimal Human Effort via Cross-Modal Synthesis
di: Yu, Luca Jiang-Tao, et al.
Pubblicazione: (2024)
di: Yu, Luca Jiang-Tao, et al.
Pubblicazione: (2024)
NeuroIncept Decoder for High-Fidelity Speech Reconstruction from Neural Activity
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
di: Khanday, Owais Mujtaba, et al.
Pubblicazione: (2025)
Sound2Hap: Learning Audio-to-Vibrotactile Haptic Generation from Human Ratings
di: Li, Yinan, et al.
Pubblicazione: (2026)
di: Li, Yinan, et al.
Pubblicazione: (2026)
A Mapping Strategy for Interacting with Latent Audio Synthesis Using Artistic Materials
di: Zheng, Shuoyang, et al.
Pubblicazione: (2024)
di: Zheng, Shuoyang, et al.
Pubblicazione: (2024)
Enhancing DMI Interactions by Integrating Haptic Feedback for Intricate Vibrato Technique
di: Piao, Ziyue, et al.
Pubblicazione: (2024)
di: Piao, Ziyue, et al.
Pubblicazione: (2024)
Towards Temporally Explainable Dysarthric Speech Clarity Assessment
di: Park, Seohyun, et al.
Pubblicazione: (2025)
di: Park, Seohyun, et al.
Pubblicazione: (2025)
A cross-talk robust multichannel VAD model for multiparty agent interactions trained using synthetic re-recordings
di: Han, Hyewon, et al.
Pubblicazione: (2024)
di: Han, Hyewon, et al.
Pubblicazione: (2024)
A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
di: Li, Yue, et al.
Pubblicazione: (2024)
di: Li, Yue, et al.
Pubblicazione: (2024)
Seeing Beyond Sound: Visualization and Abstraction in Audio Data Representation
di: Blum'e, Ashlae
Pubblicazione: (2025)
di: Blum'e, Ashlae
Pubblicazione: (2025)
Interfacing with history: Curating with audio augmented objects
di: Cliffe, Laurence
Pubblicazione: (2024)
di: Cliffe, Laurence
Pubblicazione: (2024)
Transhuman Ansambl - Voice Beyond Language
di: Ivsic, Lucija, et al.
Pubblicazione: (2024)
di: Ivsic, Lucija, et al.
Pubblicazione: (2024)
How Private is Low-Frequency Speech Audio in the Wild? An Analysis of Verbal Intelligibility by Humans and Machines
di: Liu, Ailin, et al.
Pubblicazione: (2024)
di: Liu, Ailin, et al.
Pubblicazione: (2024)
Cervical Auscultation Machine Learning for Dysphagia Assessment
di: Chia, An An, et al.
Pubblicazione: (2024)
di: Chia, An An, et al.
Pubblicazione: (2024)
ExSampling: a system for the real-time ensemble performance of field-recorded environmental sounds
di: Kobayashi, Atsuya, et al.
Pubblicazione: (2020)
di: Kobayashi, Atsuya, et al.
Pubblicazione: (2020)
Adapting Whisper for Lightweight and Efficient Automatic Speech Recognition of Children for On-device Edge Applications
di: Dutta, Satwik, et al.
Pubblicazione: (2025)
di: Dutta, Satwik, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection
di: Fan, Cunhang, et al.
Pubblicazione: (2025) -
DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
di: Yan, Sheng, et al.
Pubblicazione: (2024) -
Evaluating Spatialized Auditory Cues for Rapid Attention Capture in XR
di: Kim, Yoonsang, et al.
Pubblicazione: (2026) -
M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction
di: Fan, Cunhang, et al.
Pubblicazione: (2025) -
Auditory Attention Decoding without Spatial Information: A Diotic EEG Study
di: Yoshino, Masahiro, et al.
Pubblicazione: (2026)