Gespeichert in:
| Hauptverfasser: | Rascon, Caleb, Gato-Diaz, Luis, García-Alarcón, Eduardo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2507.02755 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Direction of Arrival Correction through Speech Quality Feedback
von: Rascon, Caleb
Veröffentlicht: (2024)
von: Rascon, Caleb
Veröffentlicht: (2024)
Scattering Transform for Auditory Attention Decoding
von: Pallenberg, René, et al.
Veröffentlicht: (2026)
von: Pallenberg, René, et al.
Veröffentlicht: (2026)
PlumberNet: Fixing interference leakage after GEV beamforming
von: Grondin, François, et al.
Veröffentlicht: (2023)
von: Grondin, François, et al.
Veröffentlicht: (2023)
Auditory Intelligence: Understanding the World Through Sound
von: Nam, Hyeonuk
Veröffentlicht: (2025)
von: Nam, Hyeonuk
Veröffentlicht: (2025)
Moravec's Paradox: Towards an Auditory Turing Test
von: Noever, David, et al.
Veröffentlicht: (2025)
von: Noever, David, et al.
Veröffentlicht: (2025)
APG-MOS: Auditory Perception Guided-MOS Predictor for Synthetic Speech
von: Lian, Zhicheng, et al.
Veröffentlicht: (2025)
von: Lian, Zhicheng, et al.
Veröffentlicht: (2025)
The MUSE Benchmark: Probing Music Perception and Auditory Relational Reasoning in Audio LLMS
von: Carone, Brandon James, et al.
Veröffentlicht: (2025)
von: Carone, Brandon James, et al.
Veröffentlicht: (2025)
A General Close-loop Predictive Coding Framework for Auditory Working Memory
von: Yuan, Zhongju, et al.
Veröffentlicht: (2025)
von: Yuan, Zhongju, et al.
Veröffentlicht: (2025)
Scaling Auditory Cognition via Test-Time Compute in Audio Language Models
von: Dang, Ting, et al.
Veröffentlicht: (2025)
von: Dang, Ting, et al.
Veröffentlicht: (2025)
AAD-LLM: Neural Attention-Driven Auditory Scene Understanding
von: Jiang, Xilin, et al.
Veröffentlicht: (2025)
von: Jiang, Xilin, et al.
Veröffentlicht: (2025)
SWIM: Short-Window CNN Integrated with Mamba for EEG-Based Auditory Spatial Attention Decoding
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
von: Zhang, Ziyang, et al.
Veröffentlicht: (2024)
DeepASA: An Object-Oriented Multi-Purpose Network for Auditory Scene Analysis
von: Lee, Dongheon, et al.
Veröffentlicht: (2025)
von: Lee, Dongheon, et al.
Veröffentlicht: (2025)
AudioScene: Integrating Object-Event Audio into 3D Scenes
von: Yuan, Shuaihang, et al.
Veröffentlicht: (2025)
von: Yuan, Shuaihang, et al.
Veröffentlicht: (2025)
CoComposer: LLM Multi-agent Collaborative Music Composition
von: Xing, Peiwen, et al.
Veröffentlicht: (2025)
von: Xing, Peiwen, et al.
Veröffentlicht: (2025)
Sound Scene Synthesis at the DCASE 2024 Challenge
von: Lagrange, Mathieu, et al.
Veröffentlicht: (2025)
von: Lagrange, Mathieu, et al.
Veröffentlicht: (2025)
Toward a Realistic Encoding Model of Auditory Affective Understanding in the Brain
von: Pan, Guandong, et al.
Veröffentlicht: (2025)
von: Pan, Guandong, et al.
Veröffentlicht: (2025)
SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
Deep Space Separable Distillation for Lightweight Acoustic Scene Classification
von: Ye, ShuQi, et al.
Veröffentlicht: (2024)
von: Ye, ShuQi, et al.
Veröffentlicht: (2024)
AudioLens: A Closer Look at Auditory Attribute Perception of Large Audio-Language Models
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
von: Yang, Chih-Kai, et al.
Veröffentlicht: (2025)
RIR-SF: Room Impulse Response Based Spatial Feature for Target Speech Recognition in Multi-Channel Multi-Speaker Scenarios
von: Shao, Yiwen, et al.
Veröffentlicht: (2023)
von: Shao, Yiwen, et al.
Veröffentlicht: (2023)
SoundCompass: Navigating Target Sound Extraction With Effective Directional Clue Integration In Complex Acoustic Scenes
von: Choi, Dayun, et al.
Veröffentlicht: (2025)
von: Choi, Dayun, et al.
Veröffentlicht: (2025)
Safe Guard: an LLM-agent for Real-time Voice-based Hate Speech Detection in Social Virtual Reality
von: Xu, Yiwen, et al.
Veröffentlicht: (2024)
von: Xu, Yiwen, et al.
Veröffentlicht: (2024)
Beyond Discrete Categories: Multi-Task Valence-Arousal Modeling for Pet Vocalization Analysis
von: Huang, Junyao, et al.
Veröffentlicht: (2025)
von: Huang, Junyao, et al.
Veröffentlicht: (2025)
Enhancing Music Genre Classification through Multi-Algorithm Analysis and User-Friendly Visualization
von: Kamuni, Navin, et al.
Veröffentlicht: (2024)
von: Kamuni, Navin, et al.
Veröffentlicht: (2024)
Revival with Voice: Multi-modal Controllable Text-to-Speech Synthesis
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
Advancing Multi-talker ASR Performance with Large Language Models
von: Shi, Mohan, et al.
Veröffentlicht: (2024)
von: Shi, Mohan, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness of Pre-Trained Audio Embeddings for Classification of Parkinson's Disease Speech Data
von: Postma, Emmy, et al.
Veröffentlicht: (2025)
von: Postma, Emmy, et al.
Veröffentlicht: (2025)
Fitting Auditory Filterbanks with Multiresolution Neural Networks
von: Lostanlen, Vincent, et al.
Veröffentlicht: (2023)
von: Lostanlen, Vincent, et al.
Veröffentlicht: (2023)
Leveraging Spatial Cues from Cochlear Implant Microphones to Efficiently Enhance Speech Separation in Real-World Listening Scenes
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025)
von: Olalere, Feyisayo, et al.
Veröffentlicht: (2025)
Explainable Deep Learning Analysis for Raga Identification in Indian Art Music
von: Singh, Parampreet, et al.
Veröffentlicht: (2024)
von: Singh, Parampreet, et al.
Veröffentlicht: (2024)
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
von: Xie, Jiamin, et al.
Veröffentlicht: (2025)
von: Xie, Jiamin, et al.
Veröffentlicht: (2025)
Dual-branch Graph Domain Adaptation for Cross-scenario Multi-modal Emotion Recognition
von: Shou, Yuntao, et al.
Veröffentlicht: (2026)
von: Shou, Yuntao, et al.
Veröffentlicht: (2026)
VNet: A GAN-based Multi-Tier Discriminator Network for Speech Synthesis Vocoders
von: Cao, Yubing, et al.
Veröffentlicht: (2024)
von: Cao, Yubing, et al.
Veröffentlicht: (2024)
A Multi-modal Approach to Dysarthria Detection and Severity Assessment Using Speech and Text Information
von: M, Anuprabha, et al.
Veröffentlicht: (2024)
von: M, Anuprabha, et al.
Veröffentlicht: (2024)
Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models
von: Kim, Jin Sob, et al.
Veröffentlicht: (2024)
von: Kim, Jin Sob, et al.
Veröffentlicht: (2024)
BSS-CFFMA: Cross-Domain Feature Fusion and Multi-Attention Speech Enhancement Network based on Self-Supervised Embedding
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2024)
von: Mattursun, Alimjan, et al.
Veröffentlicht: (2024)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
von: Berger, Clémentine, et al.
Veröffentlicht: (2025)
Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025)
MultiVerse: Efficient and Expressive Zero-Shot Multi-Task Text-to-Speech
von: Bak, Taejun, et al.
Veröffentlicht: (2024)
von: Bak, Taejun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Direction of Arrival Correction through Speech Quality Feedback
von: Rascon, Caleb
Veröffentlicht: (2024) -
Scattering Transform for Auditory Attention Decoding
von: Pallenberg, René, et al.
Veröffentlicht: (2026) -
PlumberNet: Fixing interference leakage after GEV beamforming
von: Grondin, François, et al.
Veröffentlicht: (2023) -
Auditory Intelligence: Understanding the World Through Sound
von: Nam, Hyeonuk
Veröffentlicht: (2025) -
Moravec's Paradox: Towards an Auditory Turing Test
von: Noever, David, et al.
Veröffentlicht: (2025)