Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ratnarajah, Anton, Ergezer, Mehmet, Nair, Arun, Athi, Mrudula |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
di: Neri, Michael, et al.
Pubblicazione: (2026)
di: Neri, Michael, et al.
Pubblicazione: (2026)
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024)
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024)
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
di: Ronchini, Francesca, et al.
Pubblicazione: (2025)
di: Ronchini, Francesca, et al.
Pubblicazione: (2025)
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2026)
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2026)
BRUDEX Database: Binaural Room Impulse Responses with Uniformly Distributed External Microphones
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
Acoustivision Pro: An Open-Source Interactive Platform for Room Impulse Response Analysis and Acoustic Characterization
di: Goswami, Mandip
Pubblicazione: (2026)
di: Goswami, Mandip
Pubblicazione: (2026)
Perceptual Noise-Masking with Music through Deep Spectral Envelope Shaping
di: Berger, Clémentine, et al.
Pubblicazione: (2025)
di: Berger, Clémentine, et al.
Pubblicazione: (2025)
Toward Fully-End-to-End Listened Speech Decoding from EEG Signals
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
di: Lee, Jihwan, et al.
Pubblicazione: (2024)
Joint Semantic Knowledge Distillation and Masked Acoustic Modeling for Full-band Speech Restoration with Improved Intelligibility
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
di: Liu, Xiaoyu, et al.
Pubblicazione: (2024)
AI-Generated Music Detection in Broadcast Monitoring
di: López-Ayala, David, et al.
Pubblicazione: (2026)
di: López-Ayala, David, et al.
Pubblicazione: (2026)
Comparison of Frequency-Fusion Mechanisms for Binaural Direction-of-Arrival Estimation for Multiple Speakers
di: Fejgin, Daniel, et al.
Pubblicazione: (2024)
di: Fejgin, Daniel, et al.
Pubblicazione: (2024)
Robust Generative Audio Quality Assessment: Disentangling Quality from Spurious Correlations
di: Huang, Kuan-Tang, et al.
Pubblicazione: (2026)
di: Huang, Kuan-Tang, et al.
Pubblicazione: (2026)
IS${}^3$ : Generic Impulsive--Stationary Sound Separation in Acoustic Scenes using Deep Filtering
di: Berger, Clémentine, et al.
Pubblicazione: (2025)
di: Berger, Clémentine, et al.
Pubblicazione: (2025)
Dynamic Multi-Species Bird Soundscape Generation with Acoustic Patterning and 3D Spatialization
di: Zhang, Ellie L., et al.
Pubblicazione: (2025)
di: Zhang, Ellie L., et al.
Pubblicazione: (2025)
Exploiting an External Microphone for Binaural RTF-Vector-Based Direction of Arrival Estimation for Multiple Speakers
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
di: Fejgin, Daniel, et al.
Pubblicazione: (2023)
Completing Sets of Prototype Transfer Functions for Subspace-based Direction of Arrival Estimation of Multiple Speakers
di: Fejgin, Daniel, et al.
Pubblicazione: (2025)
di: Fejgin, Daniel, et al.
Pubblicazione: (2025)
CSL-L2M: Controllable Song-Level Lyric-to-Melody Generation Based on Conditional Transformer with Fine-Grained Lyric and Musical Controls
di: Chai, Li, et al.
Pubblicazione: (2024)
di: Chai, Li, et al.
Pubblicazione: (2024)
Reverberation-based Features for Sound Event Localization and Detection with Distance Estimation
di: Berghi, Davide, et al.
Pubblicazione: (2025)
di: Berghi, Davide, et al.
Pubblicazione: (2025)
Efficient learning-based sound propagation for virtual and real-world audio processing applications
di: Ratnarajah, Anton Jeran
Pubblicazione: (2024)
di: Ratnarajah, Anton Jeran
Pubblicazione: (2024)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
di: Sato, Hiroshi, et al.
Pubblicazione: (2024)
di: Sato, Hiroshi, et al.
Pubblicazione: (2024)
Tracking of Intermittent and Moving Speakers : Dataset and Metrics
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
di: Iatariene, Taous, et al.
Pubblicazione: (2025)
Compressing Quaternion Convolutional Neural Networks for Audio Classification
di: Singh, Arshdeep, et al.
Pubblicazione: (2025)
di: Singh, Arshdeep, et al.
Pubblicazione: (2025)
Construction and Evaluation of Mandarin Multimodal Emotional Speech Database
di: Ting, Zhu, et al.
Pubblicazione: (2024)
di: Ting, Zhu, et al.
Pubblicazione: (2024)
Neural Speech and Audio Coding: Modern AI Technology Meets Traditional Codecs
di: Kim, Minje, et al.
Pubblicazione: (2024)
di: Kim, Minje, et al.
Pubblicazione: (2024)
Differentiable Modal Synthesis for Physical Modeling of Planar String Sound and Motion Simulation
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
di: Lee, Jin Woo, et al.
Pubblicazione: (2024)
Automatic Speech Recognition using Advanced Deep Learning Approaches: A survey
di: Kheddar, Hamza, et al.
Pubblicazione: (2024)
di: Kheddar, Hamza, et al.
Pubblicazione: (2024)
U-DREAM: Unsupervised Dereverberation guided by a Reverberation Model
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
Speech Enhancement Based on Drifting Models
di: Xu, Liang, et al.
Pubblicazione: (2026)
di: Xu, Liang, et al.
Pubblicazione: (2026)
A Hybrid Model for Weakly-Supervised Speech Dereverberation
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
di: Bahrman, Louis, et al.
Pubblicazione: (2025)
SWIM: Short-Window CNN Integrated with Mamba for EEG-Based Auditory Spatial Attention Decoding
di: Zhang, Ziyang, et al.
Pubblicazione: (2024)
di: Zhang, Ziyang, et al.
Pubblicazione: (2024)
Wavetable Synthesis Using CVAE for Timbre Control Based on Semantic Label
di: Yutani, Tsugumasa, et al.
Pubblicazione: (2024)
di: Yutani, Tsugumasa, et al.
Pubblicazione: (2024)
Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
di: Bae, Hanbin, et al.
Pubblicazione: (2024)
Single-stage TTS with Masked Audio Token Modeling and Semantic Knowledge Distillation
di: Gállego, Gerard I., et al.
Pubblicazione: (2024)
di: Gállego, Gerard I., et al.
Pubblicazione: (2024)
Tool Wear Prediction in CNC Turning Operations using Ultrasonic Microphone Arrays and CNNs
di: Steckel, Jan, et al.
Pubblicazione: (2024)
di: Steckel, Jan, et al.
Pubblicazione: (2024)
Classification of Heart Sounds Using Multi-Branch Deep Convolutional Network and LSTM-CNN
di: Latifi, Seyed Amir, et al.
Pubblicazione: (2024)
di: Latifi, Seyed Amir, et al.
Pubblicazione: (2024)
VoicePrompter: Robust Zero-Shot Voice Conversion with Voice Prompt and Conditional Flow Matching
di: Choi, Ha-Yeong, et al.
Pubblicazione: (2025)
di: Choi, Ha-Yeong, et al.
Pubblicazione: (2025)
A Domain-Knowledge-Inspired Music Embedding Space and a Novel Attention Mechanism for Symbolic Music Modeling
di: Guo, Z., et al.
Pubblicazione: (2022)
di: Guo, Z., et al.
Pubblicazione: (2022)
Resounding Acoustic Fields with Reciprocity
di: Lan, Zitong, et al.
Pubblicazione: (2025)
di: Lan, Zitong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Dependence on Early and Late Reverberation of Single-Channel Speaker Distance Estimation
di: Neri, Michael, et al.
Pubblicazione: (2026) -
Towards Low-Latency Tracking of Multiple Speakers With Short-Context Speaker Embeddings
di: Iatariene, Taous, et al.
Pubblicazione: (2025) -
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
di: Kim, Ji-Hoon, et al.
Pubblicazione: (2024) -
Mind the Prompt: Prompting Strategies in Audio Generations for Improving Sound Classification
di: Ronchini, Francesca, et al.
Pubblicazione: (2025) -
Velocity Potential Neural Field for Efficient Ambisonics Impulse Response Modeling
di: Masuyama, Yoshiki, et al.
Pubblicazione: (2026)