Interaural time difference loss for binaural target sound extraction
Fuente:
arXiv
Salvato in:
| Autori principali: | Hernandez-Olivan, Carlos, Delcroix, Marc, Ochiai, Tsubasa, Tawara, Naohiro, Nakatani, Tomohiro, Araki, Shoko |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024)
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024)
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
di: Tammen, Marvin, et al.
Pubblicazione: (2024)
di: Tammen, Marvin, et al.
Pubblicazione: (2024)
Mamba-based Segmentation Model for Speaker Diarization
di: Plaquet, Alexis, et al.
Pubblicazione: (2024)
di: Plaquet, Alexis, et al.
Pubblicazione: (2024)
Target Speech Extraction with Pre-trained Self-supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2024)
di: Peng, Junyi, et al.
Pubblicazione: (2024)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
di: Plaquet, Alexis, et al.
Pubblicazione: (2025)
di: Plaquet, Alexis, et al.
Pubblicazione: (2025)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025)
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025)
Reference Microphone Selection for Guided Source Separation based on the Normalized L-p Norm
di: Lohmann, Anselm, et al.
Pubblicazione: (2025)
di: Lohmann, Anselm, et al.
Pubblicazione: (2025)
Probing Self-supervised Learning Models with Target Speech Extraction
di: Peng, Junyi, et al.
Pubblicazione: (2024)
di: Peng, Junyi, et al.
Pubblicazione: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2025)
di: Peng, Junyi, et al.
Pubblicazione: (2025)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
MOVER: Combining Multiple Meeting Recognition Systems
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
Frontend Token Enhancement for Token-Based Speech Recognition
di: Ashihara, Takanori, et al.
Pubblicazione: (2026)
di: Ashihara, Takanori, et al.
Pubblicazione: (2026)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
Description and Discussion on DCASE 2025 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
di: Yasuda, Masahiro, et al.
Pubblicazione: (2025)
di: Yasuda, Masahiro, et al.
Pubblicazione: (2025)
The role of direct sound spherical harmonics representation in externalization using binaural reproduction
di: Miller, Eran, et al.
Pubblicazione: (2024)
di: Miller, Eran, et al.
Pubblicazione: (2024)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
di: Sato, Hiroshi, et al.
Pubblicazione: (2024)
di: Sato, Hiroshi, et al.
Pubblicazione: (2024)
Multizone sound field reproduction with direction-of-arrival-distribution-based regularization and its application to binaural-centered mode-matching
di: Matsuda, Ryo, et al.
Pubblicazione: (2025)
di: Matsuda, Ryo, et al.
Pubblicazione: (2025)
Investigation of Speaker Representation for Target-Speaker Speech Processing
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
Human-mimetic binaural ear design and sound source direction estimation for task realization of musculoskeletal humanoids
di: Omura, Yusuke, et al.
Pubblicazione: (2024)
di: Omura, Yusuke, et al.
Pubblicazione: (2024)
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
di: Kamo, Naoyuki, et al.
Pubblicazione: (2024)
di: Kamo, Naoyuki, et al.
Pubblicazione: (2024)
End-to-end multi-channel speaker extraction and binaural speech synthesis
di: Chi, Cheng, et al.
Pubblicazione: (2024)
di: Chi, Cheng, et al.
Pubblicazione: (2024)
Description and Discussion on DCASE 2026 Challenge Task 4: Spatial Semantic Segmentation of Sound Scenes
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2026)
di: Nguyen, Binh Thien, et al.
Pubblicazione: (2026)
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
Towards a generalized monaural and binaural auditory model for psychoacoustics and speech intelligibility
di: Biberger, Thomas, et al.
Pubblicazione: (2021)
di: Biberger, Thomas, et al.
Pubblicazione: (2021)
Real-time multichannel deep speech enhancement in hearing aids: Comparing monaural and binaural processing in complex acoustic scenarios
di: Westhausen, Nils L., et al.
Pubblicazione: (2024)
di: Westhausen, Nils L., et al.
Pubblicazione: (2024)
A tunable binaural audio telepresence system capable of balancing immersive and enhanced modes
di: Hsu, Yicheng, et al.
Pubblicazione: (2024)
di: Hsu, Yicheng, et al.
Pubblicazione: (2024)
iMagLS: Interaural Level Difference with Magnitude Least-Squares Loss for Optimized First-Order Head-Related Transfer Function
di: Berebi, Or, et al.
Pubblicazione: (2023)
di: Berebi, Or, et al.
Pubblicazione: (2023)
VBx for End-to-End Neural and Clustering-based Diarization
di: Pálka, Petr, et al.
Pubblicazione: (2025)
di: Pálka, Petr, et al.
Pubblicazione: (2025)
Onset and offset weighted loss function for sound event detection
di: Song, Tao
Pubblicazione: (2024)
di: Song, Tao
Pubblicazione: (2024)
Signal processing algorithm effective for sound quality of hearing loss simulators
di: Irino, Toshio, et al.
Pubblicazione: (2024)
di: Irino, Toshio, et al.
Pubblicazione: (2024)
Hierarchical speaker representation for target speaker extraction
di: He, Shulin, et al.
Pubblicazione: (2022)
di: He, Shulin, et al.
Pubblicazione: (2022)
Improving curriculum learning for target speaker extraction with synthetic speakers
di: Liu, Yun, et al.
Pubblicazione: (2024)
di: Liu, Yun, et al.
Pubblicazione: (2024)
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
di: Pálka, Petr, et al.
Pubblicazione: (2024)
di: Pálka, Petr, et al.
Pubblicazione: (2024)
30+ Years of Source Separation Research: Achievements and Future Challenges
di: Araki, Shoko, et al.
Pubblicazione: (2025)
di: Araki, Shoko, et al.
Pubblicazione: (2025)
Documenti analoghi
-
SoundBeam meets M2D: Target Sound Extraction with Audio Foundation Model
di: Hernandez-Olivan, Carlos, et al.
Pubblicazione: (2024) -
Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
di: Tammen, Marvin, et al.
Pubblicazione: (2024) -
Mamba-based Segmentation Model for Speaker Diarization
di: Plaquet, Alexis, et al.
Pubblicazione: (2024) -
Target Speech Extraction with Pre-trained Self-supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2024) -
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
di: Plaquet, Alexis, et al.
Pubblicazione: (2025)