An approach to optimize inference of the DIART speaker diarization pipeline
Fuente:
arXiv
Saved in:
| Main Authors: | Aperdannier, Roman, Schacht, Sigurd, Piazza, Alexander |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Review of Common Online Speaker Diarization Methods
by: Aperdannier, Roman, et al.
Published: (2024)
by: Aperdannier, Roman, et al.
Published: (2024)
Systematic Evaluation of Online Speaker Diarization Systems Regarding their Latency
by: Aperdannier, Roman, et al.
Published: (2024)
by: Aperdannier, Roman, et al.
Published: (2024)
On the calibration of powerset speaker diarization models
by: Plaquet, Alexis, et al.
Published: (2024)
by: Plaquet, Alexis, et al.
Published: (2024)
EEND-M2F: Masked-attention mask transformers for speaker diarization
by: Härkönen, Marc, et al.
Published: (2024)
by: Härkönen, Marc, et al.
Published: (2024)
LLM-based speaker diarization correction: A generalizable approach
by: Efstathiadis, Georgios, et al.
Published: (2024)
by: Efstathiadis, Georgios, et al.
Published: (2024)
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
by: Cord-Landwehr, Tobias, et al.
Published: (2025)
by: Cord-Landwehr, Tobias, et al.
Published: (2025)
SCDiar: a streaming diarization system based on speaker change detection and speech recognition
by: Zheng, Naijun, et al.
Published: (2025)
by: Zheng, Naijun, et al.
Published: (2025)
Online speaker diarization of meetings guided by speech separation
by: Gruttadauria, Elio, et al.
Published: (2024)
by: Gruttadauria, Elio, et al.
Published: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
by: Grossman, Raymond, et al.
Published: (2025)
by: Grossman, Raymond, et al.
Published: (2025)
A Benchmark for Multi-speaker Anonymization
by: Miao, Xiaoxiao, et al.
Published: (2024)
by: Miao, Xiaoxiao, et al.
Published: (2024)
Extending Whisper with prompt tuning to target-speaker ASR
by: Ma, Hao, et al.
Published: (2023)
by: Ma, Hao, et al.
Published: (2023)
You don't understand me!: Comparing ASR results for L1 and L2 speakers of Swedish
by: Cumbal, Ronald, et al.
Published: (2024)
by: Cumbal, Ronald, et al.
Published: (2024)
A multi-speaker multi-lingual voice cloning system based on vits2 for limmits 2024 challenge
by: Wang, Xiaopeng, et al.
Published: (2024)
by: Wang, Xiaopeng, et al.
Published: (2024)
Hierarchical speaker representation for target speaker extraction
by: He, Shulin, et al.
Published: (2022)
by: He, Shulin, et al.
Published: (2022)
Improving curriculum learning for target speaker extraction with synthetic speakers
by: Liu, Yun, et al.
Published: (2024)
by: Liu, Yun, et al.
Published: (2024)
Text adaptation for speaker verification with speaker-text factorized embeddings
by: Yang, Yexin, et al.
Published: (2025)
by: Yang, Yexin, et al.
Published: (2025)
Using Phonemes in cascaded S2S translation pipeline
by: Pilz, Rene, et al.
Published: (2025)
by: Pilz, Rene, et al.
Published: (2025)
VoxHakka: A Dialectally Diverse Multi-speaker Text-to-Speech System for Taiwanese Hakka
by: Chen, Li-Wei, et al.
Published: (2024)
by: Chen, Li-Wei, et al.
Published: (2024)
An efficient text augmentation approach for contextualized Mandarin speech recognition
by: Zheng, Naijun, et al.
Published: (2024)
by: Zheng, Naijun, et al.
Published: (2024)
A two-stage transliteration approach to improve performance of a multilingual ASR
by: Kumar, Rohit
Published: (2024)
by: Kumar, Rohit
Published: (2024)
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
by: Sivasankaran, Sunit, et al.
Published: (2024)
by: Sivasankaran, Sunit, et al.
Published: (2024)
Adjust-free adversarial example generation in speech recognition using evolutionary multi-objective optimization under black-box condition
by: Ishida, Shoma, et al.
Published: (2020)
by: Ishida, Shoma, et al.
Published: (2020)
Pisets: A Robust Speech Recognition System for Lectures and Interviews
by: Bondarenko, Ivan, et al.
Published: (2026)
by: Bondarenko, Ivan, et al.
Published: (2026)
An audio-quality-based multi-strategy approach for target speaker extraction in the MISP 2023 Challenge
by: Han, Runduo, et al.
Published: (2024)
by: Han, Runduo, et al.
Published: (2024)
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
by: Kopparapu, Sunil Kumar, et al.
Published: (2024)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
by: Nguyen, Thai-Binh, et al.
Published: (2024)
by: Nguyen, Thai-Binh, et al.
Published: (2024)
A dual task learning approach to fine-tune a multilingual semantic speech encoder for Spoken Language Understanding
by: Laperrière, Gaëlle, et al.
Published: (2024)
by: Laperrière, Gaëlle, et al.
Published: (2024)
Convoifilter: A case study of doing cocktail party speech recognition
by: Nguyen, Thai-Binh, et al.
Published: (2023)
by: Nguyen, Thai-Binh, et al.
Published: (2023)
X-CrossNet: A complex spectral mapping approach to target speaker extraction with cross attention speaker embedding fusion
by: Sun, Chang, et al.
Published: (2024)
by: Sun, Chang, et al.
Published: (2024)
Joint vs Sequential Speaker-Role Detection and Automatic Speech Recognition for Air-traffic Control
by: Blatt, Alexander, et al.
Published: (2024)
by: Blatt, Alexander, et al.
Published: (2024)
Towards continually learning new languages
by: Pham, Ngoc-Quan, et al.
Published: (2022)
by: Pham, Ngoc-Quan, et al.
Published: (2022)
How phonemes contribute to deep speaker models?
by: Li, Pengqi, et al.
Published: (2024)
by: Li, Pengqi, et al.
Published: (2024)
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
by: Liu, Alexander H., et al.
Published: (2024)
by: Liu, Alexander H., et al.
Published: (2024)
Accent conversion using discrete units with parallel data synthesized from controllable accented TTS
by: Nguyen, Tuan Nam, et al.
Published: (2024)
by: Nguyen, Tuan Nam, et al.
Published: (2024)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
by: Ugan, Enes Yavuz, et al.
Published: (2025)
by: Ugan, Enes Yavuz, et al.
Published: (2025)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
by: Zink, Oswald, et al.
Published: (2024)
by: Zink, Oswald, et al.
Published: (2024)
USAD: Universal Speech and Audio Representation via Distillation
by: Chang, Heng-Jui, et al.
Published: (2025)
by: Chang, Heng-Jui, et al.
Published: (2025)
Streaming Non-Autoregressive Model for Accent Conversion and Pronunciation Improvement
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
by: Nguyen, Tuan-Nam, et al.
Published: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
by: Kocour, Martin, et al.
Published: (2025)
by: Kocour, Martin, et al.
Published: (2025)
The importance of spatial and spectral information in multiple speaker tracking
by: Beit-On, Hanan, et al.
Published: (2024)
by: Beit-On, Hanan, et al.
Published: (2024)
Similar Items
-
A Review of Common Online Speaker Diarization Methods
by: Aperdannier, Roman, et al.
Published: (2024) -
Systematic Evaluation of Online Speaker Diarization Systems Regarding their Latency
by: Aperdannier, Roman, et al.
Published: (2024) -
On the calibration of powerset speaker diarization models
by: Plaquet, Alexis, et al.
Published: (2024) -
EEND-M2F: Masked-attention mask transformers for speaker diarization
by: Härkönen, Marc, et al.
Published: (2024) -
LLM-based speaker diarization correction: A generalizable approach
by: Efstathiadis, Georgios, et al.
Published: (2024)