DiariST: Streaming Speech Translation with Speaker Diarization
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Mu, Kanda, Naoyuki, Wang, Xiaofei, Chen, Junkun, Wang, Peidong, Xue, Jian, Li, Jinyu, Yoshioka, Takuya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
PHRASED: Phrase Dictionary Biasing for Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
von: Wang, Peidong, et al.
Veröffentlicht: (2025)
Profile-Error-Tolerant Target-Speaker Voice Activity Detection
von: Wang, Dongmei, et al.
Veröffentlicht: (2023)
von: Wang, Dongmei, et al.
Veröffentlicht: (2023)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)
von: Wang, Xiaofei, et al.
Veröffentlicht: (2023)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Streaming Sortformer: Speaker Cache-Based Online Speaker Diarization with Arrival-Time Ordering
von: Medennikov, Ivan, et al.
Veröffentlicht: (2025)
von: Medennikov, Ivan, et al.
Veröffentlicht: (2025)
DiariZen Explained: A Tutorial for the Open Source State-of-the-Art Speaker Diarization Pipeline
von: Raghav, Nikhil
Veröffentlicht: (2026)
von: Raghav, Nikhil
Veröffentlicht: (2026)
Multichannel Long-Term Streaming Neural Speech Enhancement for Static and Moving Speakers
von: Quan, Changsheng, et al.
Veröffentlicht: (2024)
von: Quan, Changsheng, et al.
Veröffentlicht: (2024)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
A Review of Common Online Speaker Diarization Methods
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
Unifying Diarization, Separation, and ASR with Multi-Speaker Encoder
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2025)
System Description for the Displace Speaker Diarization Challenge 2023
von: Aliyev, Ali
Veröffentlicht: (2024)
von: Aliyev, Ali
Veröffentlicht: (2024)
Improving Speaker Diarization using Semantic Information: Joint Pairwise Constraints Propagation
von: Cheng, Luyao, et al.
Veröffentlicht: (2023)
von: Cheng, Luyao, et al.
Veröffentlicht: (2023)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
USED: Universal Speaker Extraction and Diarization
von: Ao, Junyi, et al.
Veröffentlicht: (2023)
von: Ao, Junyi, et al.
Veröffentlicht: (2023)
Systematic Evaluation of Online Speaker Diarization Systems Regarding their Latency
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
von: Aperdannier, Roman, et al.
Veröffentlicht: (2024)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
von: Pálka, Petr, et al.
Veröffentlicht: (2024)
LS-EEND: Long-Form Streaming End-to-End Neural Diarization with Online Attractor Extraction
von: Liang, Di, et al.
Veröffentlicht: (2024)
von: Liang, Di, et al.
Veröffentlicht: (2024)
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
von: Ghosh, Sreyan, et al.
Veröffentlicht: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
von: Dai, Yuhang, et al.
Veröffentlicht: (2025)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
von: Jin, Zengrui, et al.
Veröffentlicht: (2024)
StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
von: Guo, Shoutao, et al.
Veröffentlicht: (2025)
DiCoW: Diarization-Conditioned Whisper for Target Speaker Automatic Speech Recognition
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
von: Polok, Alexander, et al.
Veröffentlicht: (2024)
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
Leveraging Self-Supervised Learning for Speaker Diarization
von: Han, Jiangyu, et al.
Veröffentlicht: (2024)
von: Han, Jiangyu, et al.
Veröffentlicht: (2024)
Mamba-based Segmentation Model for Speaker Diarization
von: Plaquet, Alexis, et al.
Veröffentlicht: (2024)
von: Plaquet, Alexis, et al.
Veröffentlicht: (2024)
TransVIP: Speech to Speech Translation System with Voice and Isochrony Preservation
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
von: Le, Chenyang, et al.
Veröffentlicht: (2024)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
von: Horiguchi, Shota, et al.
Veröffentlicht: (2025)
Speaker Embeddings With Weakly Supervised Voice Activity Detection For Efficient Speaker Diarization
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
von: Thienpondt, Jenthe, et al.
Veröffentlicht: (2024)
Length Aware Speech Translation for Video Dubbing
von: Chadha, Harveen Singh, et al.
Veröffentlicht: (2025)
von: Chadha, Harveen Singh, et al.
Veröffentlicht: (2025)
Two-stage Framework for Robust Speech Emotion Recognition Using Target Speaker Extraction in Human Speech Noise Conditions
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
von: Mi, Jinyi, et al.
Veröffentlicht: (2024)
Investigation of Speaker Representation for Target-Speaker Speech Processing
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Uncertainty Quantification in Machine Learning for Joint Speaker Diarization and Identification
von: McKnight, Simon W., et al.
Veröffentlicht: (2023)
von: McKnight, Simon W., et al.
Veröffentlicht: (2023)
StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
DiarizationLM: Speaker Diarization Post-Processing with Large Language Models
von: Wang, Quan, et al.
Veröffentlicht: (2024)
von: Wang, Quan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2025) -
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2024) -
PHRASED: Phrase Dictionary Biasing for Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2025) -
Profile-Error-Tolerant Target-Speaker Voice Activity Detection
von: Wang, Dongmei, et al.
Veröffentlicht: (2023) -
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)