Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
Fuente:
arXiv
Salvato in:
| Autori principali: | Boeddeker, Christoph, Cord-Landwehr, Tobias, Haeb-Umbach, Reinhold |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024)
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2025)
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2025)
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments
di: Meise, Adrian, et al.
Pubblicazione: (2026)
di: Meise, Adrian, et al.
Pubblicazione: (2026)
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
di: Kuhlmann, Michael, et al.
Pubblicazione: (2026)
di: Kuhlmann, Michael, et al.
Pubblicazione: (2026)
On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation
di: Meise, Adrian, et al.
Pubblicazione: (2025)
di: Meise, Adrian, et al.
Pubblicazione: (2025)
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
di: Deegen, Marc, et al.
Pubblicazione: (2026)
di: Deegen, Marc, et al.
Pubblicazione: (2026)
TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
di: Boeddeker, Christoph, et al.
Pubblicazione: (2023)
di: Boeddeker, Christoph, et al.
Pubblicazione: (2023)
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
di: von Neumann, Thilo, et al.
Pubblicazione: (2025)
di: von Neumann, Thilo, et al.
Pubblicazione: (2025)
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
di: von Neumann, Thilo, et al.
Pubblicazione: (2023)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025)
di: Haeb-Umbach, Reinhold, et al.
Pubblicazione: (2025)
Towards Frame-level Quality Predictions of Synthetic Speech
di: Kuhlmann, Michael, et al.
Pubblicazione: (2025)
di: Kuhlmann, Michael, et al.
Pubblicazione: (2025)
Diminishing Domain Mismatch for DNN-Based Acoustic Distance Estimation via Stochastic Room Reverberation Models
di: Gburrek, Tobias, et al.
Pubblicazione: (2024)
di: Gburrek, Tobias, et al.
Pubblicazione: (2024)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
di: Rautenberg, Frederik, et al.
Pubblicazione: (2026)
di: Rautenberg, Frederik, et al.
Pubblicazione: (2026)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
di: Kuhlmann, Michael, et al.
Pubblicazione: (2026)
di: Kuhlmann, Michael, et al.
Pubblicazione: (2026)
Error Analysis in a Modular Meeting Transcription System
di: Vieting, Peter, et al.
Pubblicazione: (2025)
di: Vieting, Peter, et al.
Pubblicazione: (2025)
Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
di: Vieting, Peter, et al.
Pubblicazione: (2023)
di: Vieting, Peter, et al.
Pubblicazione: (2023)
Synthesizing speech with selected perceptual voice qualities - A case study with creaky voice
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
Speaker and Style Disentanglement of Speech Based on Contrastive Predictive Coding Supported Factorized Variational Autoencoder
di: Xie, Yuying, et al.
Pubblicazione: (2024)
di: Xie, Yuying, et al.
Pubblicazione: (2024)
Speech Synthesis along Perceptual Voice Quality Dimensions
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
di: Rautenberg, Frederik, et al.
Pubblicazione: (2025)
30+ Years of Source Separation Research: Achievements and Future Challenges
di: Araki, Shoko, et al.
Pubblicazione: (2025)
di: Araki, Shoko, et al.
Pubblicazione: (2025)
DM-ASR: Diarization-aware Multi-speaker ASR with Large Language Models
di: Li, Li, et al.
Pubblicazione: (2026)
di: Li, Li, et al.
Pubblicazione: (2026)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
di: Cornell, Samuele, et al.
Pubblicazione: (2024)
PixIT: Joint Training of Speaker Diarization and Speech Separation from Real-world Multi-speaker Recordings
di: Kalda, Joonas, et al.
Pubblicazione: (2024)
di: Kalda, Joonas, et al.
Pubblicazione: (2024)
Improving curriculum learning for target speaker extraction with synthetic speakers
di: Liu, Yun, et al.
Pubblicazione: (2024)
di: Liu, Yun, et al.
Pubblicazione: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
di: Grossman, Raymond, et al.
Pubblicazione: (2025)
di: Grossman, Raymond, et al.
Pubblicazione: (2025)
Online speaker diarization of meetings guided by speech separation
di: Gruttadauria, Elio, et al.
Pubblicazione: (2024)
di: Gruttadauria, Elio, et al.
Pubblicazione: (2024)
Hierarchical speaker representation for target speaker extraction
di: He, Shulin, et al.
Pubblicazione: (2022)
di: He, Shulin, et al.
Pubblicazione: (2022)
Triage knowledge distillation for speaker verification
di: Kim, Ju-ho, et al.
Pubblicazione: (2026)
di: Kim, Ju-ho, et al.
Pubblicazione: (2026)
Privacy-oriented manipulation of speaker representations
di: Teixeira, Francisco, et al.
Pubblicazione: (2023)
di: Teixeira, Francisco, et al.
Pubblicazione: (2023)
Text adaptation for speaker verification with speaker-text factorized embeddings
di: Yang, Yexin, et al.
Pubblicazione: (2025)
di: Yang, Yexin, et al.
Pubblicazione: (2025)
Non-locally averaged pruned reassigned spectrograms: a tool for glottal pulse visualization and analysis
di: Griswold, Gabriel J., et al.
Pubblicazione: (2025)
di: Griswold, Gabriel J., et al.
Pubblicazione: (2025)
Improving fairness in speaker verification via Group-adapted Fusion Network
di: Shen, Hua, et al.
Pubblicazione: (2022)
di: Shen, Hua, et al.
Pubblicazione: (2022)
VBx for End-to-End Neural and Clustering-based Diarization
di: Pálka, Petr, et al.
Pubblicazione: (2025)
di: Pálka, Petr, et al.
Pubblicazione: (2025)
Spatially-Augmented Sequence-to-Sequence Neural Diarization for Meetings
di: Li, Li, et al.
Pubblicazione: (2025)
di: Li, Li, et al.
Pubblicazione: (2025)
Token-based Attractors and Cross-attention in Spoof Diarization
di: Koo, Kyo-Won, et al.
Pubblicazione: (2025)
di: Koo, Kyo-Won, et al.
Pubblicazione: (2025)
Curriculum learning for self-supervised speaker verification
di: Heo, Hee-Soo, et al.
Pubblicazione: (2022)
di: Heo, Hee-Soo, et al.
Pubblicazione: (2022)
Improving Neural Diarization through Speaker Attribute Attractors and Local Dependency Modeling
di: Palzer, David, et al.
Pubblicazione: (2025)
di: Palzer, David, et al.
Pubblicazione: (2025)
Prompt-driven Target Speech Diarization
di: Jiang, Yidi, et al.
Pubblicazione: (2023)
di: Jiang, Yidi, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024) -
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2024) -
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
di: Cord-Landwehr, Tobias, et al.
Pubblicazione: (2025) -
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
di: von Neumann, Thilo, et al.
Pubblicazione: (2023) -
Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments
di: Meise, Adrian, et al.
Pubblicazione: (2026)