Alignment-Free Training for Transducer-based Multi-Talker ASR
Fuente:
arXiv
Salvato in:
| Autori principali: | Moriya, Takafumi, Horiguchi, Shota, Delcroix, Marc, Masumura, Ryo, Ashihara, Takanori, Sato, Hiroshi, Matsuura, Kohei, Mimura, Masato |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
di: Moriya, Takafumi, et al.
Pubblicazione: (2024)
Generic Speech Enhancement with Self-Supervised Representation Space Loss
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
di: Sato, Hiroshi, et al.
Pubblicazione: (2025)
Investigation of Speaker Representation for Target-Speaker Speech Processing
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)
SpeakerBeam-SS: Real-time Target Speaker Extraction with Lightweight Conv-TasNet and State Space Modeling
di: Sato, Hiroshi, et al.
Pubblicazione: (2024)
di: Sato, Hiroshi, et al.
Pubblicazione: (2024)
Frontend Token Enhancement for Token-Based Speech Recognition
di: Ashihara, Takanori, et al.
Pubblicazione: (2026)
di: Ashihara, Takanori, et al.
Pubblicazione: (2026)
Can We Really Repurpose Multi-Speaker ASR Corpus for Speaker Diarization?
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
di: Moriya, Takafumi, et al.
Pubblicazione: (2025)
di: Moriya, Takafumi, et al.
Pubblicazione: (2025)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
di: Ashihara, Takanori, et al.
Pubblicazione: (2024)
Factor-Conditioned Speaking-Style Captioning
di: Ando, Atsushi, et al.
Pubblicazione: (2024)
di: Ando, Atsushi, et al.
Pubblicazione: (2024)
Mitigating Non-Target Speaker Bias in Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
di: Ashihara, Takanori, et al.
Pubblicazione: (2023)
di: Ashihara, Takanori, et al.
Pubblicazione: (2023)
Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
di: Matsuura, Kohei, et al.
Pubblicazione: (2024)
di: Matsuura, Kohei, et al.
Pubblicazione: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Microphone Array Geometry Independent Multi-Talker Distant ASR: NTT System for the DASR Task of the CHiME-8 Challenge
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
di: Kamo, Naoyuki, et al.
Pubblicazione: (2025)
NTT Multi-Speaker ASR System for the DASR Task of CHiME-8 Challenge
di: Kamo, Naoyuki, et al.
Pubblicazione: (2024)
di: Kamo, Naoyuki, et al.
Pubblicazione: (2024)
Pretraining Multi-Speaker Identification for Neural Speaker Diarization
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
di: Horiguchi, Shota, et al.
Pubblicazione: (2025)
Applying LLMs for Rescoring N-best ASR Hypotheses of Casual Conversations: Effects of Domain Adaptation and Context Carry-over
di: Ogawa, Atsunori, et al.
Pubblicazione: (2024)
di: Ogawa, Atsunori, et al.
Pubblicazione: (2024)
Lightweight Zero-shot Text-to-Speech with Mixture of Adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
Probing Self-supervised Learning Models with Target Speech Extraction
di: Peng, Junyi, et al.
Pubblicazione: (2024)
di: Peng, Junyi, et al.
Pubblicazione: (2024)
Mamba-based Segmentation Model for Speaker Diarization
di: Plaquet, Alexis, et al.
Pubblicazione: (2024)
di: Plaquet, Alexis, et al.
Pubblicazione: (2024)
Dissecting the Segmentation Model of End-to-End Diarization with Vector Clustering
di: Plaquet, Alexis, et al.
Pubblicazione: (2025)
di: Plaquet, Alexis, et al.
Pubblicazione: (2025)
SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
di: Fan, Zhiyun, et al.
Pubblicazione: (2024)
di: Fan, Zhiyun, et al.
Pubblicazione: (2024)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2025)
di: Peng, Junyi, et al.
Pubblicazione: (2025)
Decoder-only Conformer with Modality-aware Sparse Mixtures of Experts for ASR
di: Lee, Jaeyoung, et al.
Pubblicazione: (2026)
di: Lee, Jaeyoung, et al.
Pubblicazione: (2026)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
di: He, Xiluo, et al.
Pubblicazione: (2025)
di: He, Xiluo, et al.
Pubblicazione: (2025)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
di: Yang, Yufeng, et al.
Pubblicazione: (2025)
di: Yang, Yufeng, et al.
Pubblicazione: (2025)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
di: Zhao, Wenbo, et al.
Pubblicazione: (2024)
di: Zhao, Wenbo, et al.
Pubblicazione: (2024)
Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition Performance
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
di: Ochiai, Tsubasa, et al.
Pubblicazione: (2024)
Promptformer: Prompted Conformer Transducer for ASR
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
di: Duarte-Torres, Sergio, et al.
Pubblicazione: (2024)
Chunkwise Aligners for Streaming Speech Recognition
di: Teo, Wen Shen, et al.
Pubblicazione: (2026)
di: Teo, Wen Shen, et al.
Pubblicazione: (2026)
Investigating the Effect of Label Topology and Training Criterion on ASR Performance and Alignment Quality
di: Raissi, Tina, et al.
Pubblicazione: (2024)
di: Raissi, Tina, et al.
Pubblicazione: (2024)
Joint Training of Speaker Embedding Extractor, Speech and Overlap Detection for Diarization
di: Pálka, Petr, et al.
Pubblicazione: (2024)
di: Pálka, Petr, et al.
Pubblicazione: (2024)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
di: Thorbecke, Iuliia, et al.
Pubblicazione: (2024)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
di: Kumar, Shashi, et al.
Pubblicazione: (2024)
Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation
di: Wang, Peidong, et al.
Pubblicazione: (2025)
di: Wang, Peidong, et al.
Pubblicazione: (2025)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
Joint ASR and Speaker Role Tagging with Serialized Output Training
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
di: Xu, Anfeng, et al.
Pubblicazione: (2025)
Training-Free Intelligibility-Guided Observation Addition for Noisy ASR
di: Li, Haoyang, et al.
Pubblicazione: (2026)
di: Li, Haoyang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Boosting Hybrid Autoregressive Transducer-based ASR with Internal Acoustic Model Training and Dual Blank Thresholding
di: Moriya, Takafumi, et al.
Pubblicazione: (2024) -
Generic Speech Enhancement with Self-Supervised Representation Space Loss
di: Sato, Hiroshi, et al.
Pubblicazione: (2025) -
Investigation of Speaker Representation for Target-Speaker Speech Processing
di: Ashihara, Takanori, et al.
Pubblicazione: (2024) -
Recursive Attentive Pooling for Extracting Speaker Embeddings from Multi-Speaker Recordings
di: Horiguchi, Shota, et al.
Pubblicazione: (2024) -
Guided Speaker Embedding
di: Horiguchi, Shota, et al.
Pubblicazione: (2024)