On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
Fuente:
arXiv
Guardado en:
| Autores principales: | Deegen, Marc, Gburrek, Tobias, Cord-Landwehr, Tobias, von Neumann, Thilo, Han, Jiangyu, Burget, Lukáš, Haeb-Umbach, Reinhold |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
por: Cord-Landwehr, Tobias, et al.
Publicado: (2025)
por: Cord-Landwehr, Tobias, et al.
Publicado: (2025)
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
por: von Neumann, Thilo, et al.
Publicado: (2023)
por: von Neumann, Thilo, et al.
Publicado: (2023)
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
por: Cord-Landwehr, Tobias, et al.
Publicado: (2024)
por: Cord-Landwehr, Tobias, et al.
Publicado: (2024)
Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
por: Boeddeker, Christoph, et al.
Publicado: (2024)
por: Boeddeker, Christoph, et al.
Publicado: (2024)
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
por: Kuhlmann, Michael, et al.
Publicado: (2026)
por: Kuhlmann, Michael, et al.
Publicado: (2026)
On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation
por: Meise, Adrian, et al.
Publicado: (2025)
por: Meise, Adrian, et al.
Publicado: (2025)
Diminishing Domain Mismatch for DNN-Based Acoustic Distance Estimation via Stochastic Room Reverberation Models
por: Gburrek, Tobias, et al.
Publicado: (2024)
por: Gburrek, Tobias, et al.
Publicado: (2024)
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
por: Cord-Landwehr, Tobias, et al.
Publicado: (2024)
por: Cord-Landwehr, Tobias, et al.
Publicado: (2024)
Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments
por: Meise, Adrian, et al.
Publicado: (2026)
por: Meise, Adrian, et al.
Publicado: (2026)
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
por: von Neumann, Thilo, et al.
Publicado: (2025)
por: von Neumann, Thilo, et al.
Publicado: (2025)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
por: Kuhlmann, Michael, et al.
Publicado: (2026)
por: Kuhlmann, Michael, et al.
Publicado: (2026)
Spatially Aware Self-Supervised Models for Multi-Channel Neural Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2025)
por: Han, Jiangyu, et al.
Publicado: (2025)
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
por: von Neumann, Thilo, et al.
Publicado: (2023)
por: von Neumann, Thilo, et al.
Publicado: (2023)
Leveraging Self-Supervised Learning for Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2024)
por: Han, Jiangyu, et al.
Publicado: (2024)
TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
por: Boeddeker, Christoph, et al.
Publicado: (2023)
por: Boeddeker, Christoph, et al.
Publicado: (2023)
Efficient and Generalizable Speaker Diarization via Structured Pruning of Self-Supervised Models
por: Han, Jiangyu, et al.
Publicado: (2025)
por: Han, Jiangyu, et al.
Publicado: (2025)
VBx for End-to-End Neural and Clustering-based Diarization
por: Pálka, Petr, et al.
Publicado: (2025)
por: Pálka, Petr, et al.
Publicado: (2025)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
por: Rautenberg, Frederik, et al.
Publicado: (2026)
por: Rautenberg, Frederik, et al.
Publicado: (2026)
Fine-tune Before Structured Pruning: Towards Compact and Accurate Self-Supervised Models for Speaker Diarization
por: Han, Jiangyu, et al.
Publicado: (2025)
por: Han, Jiangyu, et al.
Publicado: (2025)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
por: Xie, Yuying, et al.
Publicado: (2024)
por: Xie, Yuying, et al.
Publicado: (2024)
Error Analysis in a Modular Meeting Transcription System
por: Vieting, Peter, et al.
Publicado: (2025)
por: Vieting, Peter, et al.
Publicado: (2025)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
por: Vieting, Peter, et al.
Publicado: (2023)
por: Vieting, Peter, et al.
Publicado: (2023)
Towards Frame-level Quality Predictions of Synthetic Speech
por: Kuhlmann, Michael, et al.
Publicado: (2025)
por: Kuhlmann, Michael, et al.
Publicado: (2025)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
por: Pálka, Petr, et al.
Publicado: (2024)
por: Pálka, Petr, et al.
Publicado: (2024)
Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization
por: Polok, Alexander, et al.
Publicado: (2026)
por: Polok, Alexander, et al.
Publicado: (2026)
Do End-to-End Neural Diarization Attractors Need to Encode Speaker Characteristic Information?
por: Zhang, Lin, et al.
Publicado: (2024)
por: Zhang, Lin, et al.
Publicado: (2024)
DiaPer: End-to-End Neural Diarization with Perceiver-Based Attractors
por: Landini, Federico, et al.
Publicado: (2023)
por: Landini, Federico, et al.
Publicado: (2023)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
por: Haeb-Umbach, Reinhold, et al.
Publicado: (2025)
por: Haeb-Umbach, Reinhold, et al.
Publicado: (2025)
Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
por: Rautenberg, Frederik, et al.
Publicado: (2025)
por: Rautenberg, Frederik, et al.
Publicado: (2025)
Speech Synthesis along Perceptual Voice Quality Dimensions
por: Rautenberg, Frederik, et al.
Publicado: (2025)
por: Rautenberg, Frederik, et al.
Publicado: (2025)
BUT System for the MLC-SLM Challenge
por: Polok, Alexander, et al.
Publicado: (2025)
por: Polok, Alexander, et al.
Publicado: (2025)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
por: Horiguchi, Shota, et al.
Publicado: (2025)
por: Horiguchi, Shota, et al.
Publicado: (2025)
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan
por: Xu, Anfeng, et al.
Publicado: (2026)
por: Xu, Anfeng, et al.
Publicado: (2026)
Mamba-based Segmentation Model for Speaker Diarization
por: Plaquet, Alexis, et al.
Publicado: (2024)
por: Plaquet, Alexis, et al.
Publicado: (2024)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
por: Horiguchi, Shota, et al.
Publicado: (2025)
por: Horiguchi, Shota, et al.
Publicado: (2025)
30+ Years of Source Separation Research: Achievements and Future Challenges
por: Araki, Shoko, et al.
Publicado: (2025)
por: Araki, Shoko, et al.
Publicado: (2025)
Target Speaker ASR with Whisper
por: Polok, Alexander, et al.
Publicado: (2024)
por: Polok, Alexander, et al.
Publicado: (2024)
ASR-Synchronized Speaker-Role Diarization
por: Ghosh, Arindam, et al.
Publicado: (2025)
por: Ghosh, Arindam, et al.
Publicado: (2025)
SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
por: Polok, Alexander, et al.
Publicado: (2026)
por: Polok, Alexander, et al.
Publicado: (2026)
Ejemplares similares
-
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
por: Cord-Landwehr, Tobias, et al.
Publicado: (2025) -
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
por: von Neumann, Thilo, et al.
Publicado: (2023) -
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
por: Cord-Landwehr, Tobias, et al.
Publicado: (2024) -
Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
por: Boeddeker, Christoph, et al.
Publicado: (2024) -
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
por: Kuhlmann, Michael, et al.
Publicado: (2026)