Combining TF-GridNet and Mixture Encoder for Continuous Speech Separation for Meeting Transcription
Fuente:
arXiv
Saved in:
| Main Authors: | Vieting, Peter, Berger, Simon, von Neumann, Thilo, Boeddeker, Christoph, Schlüter, Ralf, Haeb-Umbach, Reinhold |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Error Analysis in a Modular Meeting Transcription System
by: Vieting, Peter, et al.
Published: (2025)
by: Vieting, Peter, et al.
Published: (2025)
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
by: von Neumann, Thilo, et al.
Published: (2023)
by: von Neumann, Thilo, et al.
Published: (2023)
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
by: von Neumann, Thilo, et al.
Published: (2023)
by: von Neumann, Thilo, et al.
Published: (2023)
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
by: Cord-Landwehr, Tobias, et al.
Published: (2024)
by: Cord-Landwehr, Tobias, et al.
Published: (2024)
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
by: von Neumann, Thilo, et al.
Published: (2025)
by: von Neumann, Thilo, et al.
Published: (2025)
TS-SEP: Joint Diarization and Separation Conditioned on Estimated Speaker Embeddings
by: Boeddeker, Christoph, et al.
Published: (2023)
by: Boeddeker, Christoph, et al.
Published: (2023)
Microphone Array Signal Processing and Deep Learning for Speech Enhancement
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
by: Haeb-Umbach, Reinhold, et al.
Published: (2025)
Speech Quality-Based Localization of Low-Quality Speech and Text-to-Speech Synthesis Artefacts
by: Kuhlmann, Michael, et al.
Published: (2026)
by: Kuhlmann, Michael, et al.
Published: (2026)
Diminishing Domain Mismatch for DNN-Based Acoustic Distance Estimation via Stochastic Room Reverberation Models
by: Gburrek, Tobias, et al.
Published: (2024)
by: Gburrek, Tobias, et al.
Published: (2024)
Speech Synthesis along Perceptual Voice Quality Dimensions
by: Rautenberg, Frederik, et al.
Published: (2025)
by: Rautenberg, Frederik, et al.
Published: (2025)
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
by: Cord-Landwehr, Tobias, et al.
Published: (2025)
by: Cord-Landwehr, Tobias, et al.
Published: (2025)
Once more Diarization: Improving meeting transcription systems through segment-level speaker reassignment
by: Boeddeker, Christoph, et al.
Published: (2024)
by: Boeddeker, Christoph, et al.
Published: (2024)
30+ Years of Source Separation Research: Achievements and Future Challenges
by: Araki, Shoko, et al.
Published: (2025)
by: Araki, Shoko, et al.
Published: (2025)
TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
by: Shin, Ui-Hyeop, et al.
Published: (2025)
by: Shin, Ui-Hyeop, et al.
Published: (2025)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
by: Vieting, Peter, et al.
Published: (2025)
by: Vieting, Peter, et al.
Published: (2025)
Speech Quality Embeddings for Improved Detection and Classification of Degradations in Speech Signals
by: Kuhlmann, Michael, et al.
Published: (2026)
by: Kuhlmann, Michael, et al.
Published: (2026)
Geodesic interpolation of frame-wise speaker embeddings for the diarization of meeting scenarios
by: Cord-Landwehr, Tobias, et al.
Published: (2024)
by: Cord-Landwehr, Tobias, et al.
Published: (2024)
Unified Learnable 2D Convolutional Feature Extraction for ASR
by: Vieting, Peter, et al.
Published: (2025)
by: Vieting, Peter, et al.
Published: (2025)
Chunked Attention-based Encoder-Decoder Model for Streaming Speech Recognition
by: Zeineldeen, Mohammad, et al.
Published: (2023)
by: Zeineldeen, Mohammad, et al.
Published: (2023)
TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement
by: Saijo, Kohei, et al.
Published: (2024)
by: Saijo, Kohei, et al.
Published: (2024)
Towards Frame-level Quality Predictions of Synthetic Speech
by: Kuhlmann, Michael, et al.
Published: (2025)
by: Kuhlmann, Michael, et al.
Published: (2025)
Disentangling Pitch and Creak for Speaker Identity Preservation in Speech Synthesis
by: Rautenberg, Frederik, et al.
Published: (2026)
by: Rautenberg, Frederik, et al.
Published: (2026)
Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments
by: Meise, Adrian, et al.
Published: (2026)
by: Meise, Adrian, et al.
Published: (2026)
Asymmetric Encoder-Decoder Based on Time-Frequency Correlation for Speech Separation
by: Shin, Ui-Hyeop, et al.
Published: (2026)
by: Shin, Ui-Hyeop, et al.
Published: (2026)
Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders
by: Shan, Weiqiao, et al.
Published: (2025)
by: Shan, Weiqiao, et al.
Published: (2025)
On the Role of Spatial Features in Foundation-Model-Based Speaker Diarization
by: Deegen, Marc, et al.
Published: (2026)
by: Deegen, Marc, et al.
Published: (2026)
TF-MLPNet: Tiny Real-Time Neural Speech Separation
by: Itani, Malek, et al.
Published: (2025)
by: Itani, Malek, et al.
Published: (2025)
On the Application of Diffusion Models for Simultaneous Denoising and Dereverberation
by: Meise, Adrian, et al.
Published: (2025)
by: Meise, Adrian, et al.
Published: (2025)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
by: Raissi, Tina, et al.
Published: (2024)
by: Raissi, Tina, et al.
Published: (2024)
The Conformer Encoder May Reverse the Time Dimension
by: Schmitt, Robin, et al.
Published: (2024)
by: Schmitt, Robin, et al.
Published: (2024)
Enhancing Lyrics Transcription on Music Mixtures with Consistency Loss
by: Huang, Jiawen, et al.
Published: (2025)
by: Huang, Jiawen, et al.
Published: (2025)
Prompting Whisper for Joint Speech Transcription and Diarization
by: Zamyrova, Mariia, et al.
Published: (2026)
by: Zamyrova, Mariia, et al.
Published: (2026)
Mel-RoFormer for Vocal Separation and Vocal Melody Transcription
by: Wang, Ju-Chiang, et al.
Published: (2024)
by: Wang, Ju-Chiang, et al.
Published: (2024)
Exploiting Music Source Separation for Automatic Lyrics Transcription with Whisper
by: Syed, Jaza, et al.
Published: (2025)
by: Syed, Jaza, et al.
Published: (2025)
The CHiME-8 DASR Challenge for Generalizable and Array Agnostic Distant Automatic Speech Recognition and Diarization
by: Cornell, Samuele, et al.
Published: (2024)
by: Cornell, Samuele, et al.
Published: (2024)
Cross-Talk Speech Reduction, by Separation, for Separation
by: Wang, Zhong-Qiu, et al.
Published: (2026)
by: Wang, Zhong-Qiu, et al.
Published: (2026)
Right Label Context in End-to-End Training of Time-Synchronous ASR Models
by: Raissi, Tina, et al.
Published: (2025)
by: Raissi, Tina, et al.
Published: (2025)
Enhanced Automatic Drum Transcription via Drum Stem Source Separation
by: Riley, Xavier, et al.
Published: (2025)
by: Riley, Xavier, et al.
Published: (2025)
Electrolaryngeal Speech Intelligibility Enhancement Through Robust Linguistic Encoders
by: Violeta, Lester Phillip, et al.
Published: (2023)
by: Violeta, Lester Phillip, et al.
Published: (2023)
MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder
by: Zhang, Bowen, et al.
Published: (2025)
by: Zhang, Bowen, et al.
Published: (2025)
Similar Items
-
Error Analysis in a Modular Meeting Transcription System
by: Vieting, Peter, et al.
Published: (2025) -
Meeting Recognition with Continuous Speech Separation and Transcription-Supported Diarization
by: von Neumann, Thilo, et al.
Published: (2023) -
MeetEval: A Toolkit for Computation of Word Error Rates for Meeting Transcription Systems
by: von Neumann, Thilo, et al.
Published: (2023) -
Simultaneous Diarization and Separation of Meetings through the Integration of Statistical Mixture Models
by: Cord-Landwehr, Tobias, et al.
Published: (2024) -
Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition
by: von Neumann, Thilo, et al.
Published: (2025)