Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
Fuente:
arXiv
Salvato in:
| Autori principali: | Kashiwagi, Yosuke, Futami, Hayato, Tsunoo, Emiru, Arora, Siddhant, Watanabe, Shinji |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024)
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024)
Differentiable K-means for Fully-optimized Discrete Token-based ASR
di: Onda, Kentaro, et al.
Pubblicazione: (2025)
di: Onda, Kentaro, et al.
Pubblicazione: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)
di: Futami, Hayato, et al.
Pubblicazione: (2025)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Finding Task-specific Subnetworks in Multi-task Spoken Language Understanding Model
di: Futami, Hayato, et al.
Pubblicazione: (2024)
di: Futami, Hayato, et al.
Pubblicazione: (2024)
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
UniverSLU: Universal Spoken Language Understanding for Diverse Tasks with Natural Language Instructions
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2025)
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2025)
Elevating Robust Multi-Talker ASR by Decoupling Speaker Separation and Speech Recognition
di: Yang, Yufeng, et al.
Pubblicazione: (2025)
di: Yang, Yufeng, et al.
Pubblicazione: (2025)
Causal Speech Enhancement with Predicting Semantics based on Quantized Self-supervised Learning Features
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
di: Tsunoo, Emiru, et al.
Pubblicazione: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
LibriTTS-VI: A Public Corpus and Novel Methods for Efficient Voice Impression Control
di: Ohmura, Junki, et al.
Pubblicazione: (2025)
di: Ohmura, Junki, et al.
Pubblicazione: (2025)
ARECHO: Autoregressive Evaluation via Chain-Based Hypothesis Optimization for Speech Multi-Metric Estimation
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
di: Shi, Jiatong, et al.
Pubblicazione: (2025)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
di: Sakuma, Asahi, et al.
Pubblicazione: (2025)
di: Sakuma, Asahi, et al.
Pubblicazione: (2025)
Phonological Tokenizer: Prosody-Aware Phonetic Token via Multi-Objective Fine-Tuning with Differentiable K-Means
di: Onda, Kentaro, et al.
Pubblicazione: (2026)
di: Onda, Kentaro, et al.
Pubblicazione: (2026)
Cross-Speaker Encoding Network for Multi-Talker Speech Recognition
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
di: Kang, Jiawen, et al.
Pubblicazione: (2024)
DialogueSidon: Recovering Full-Duplex Dialogue Tracks from In-the-Wild Dialogue Audio
di: Nakata, Wataru, et al.
Pubblicazione: (2026)
di: Nakata, Wataru, et al.
Pubblicazione: (2026)
Scaling Multi-Talker ASR with Speaker-Agnostic Activity Streams
di: He, Xiluo, et al.
Pubblicazione: (2025)
di: He, Xiluo, et al.
Pubblicazione: (2025)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
di: Aldeneh, Zakaria, et al.
Pubblicazione: (2024)
di: Aldeneh, Zakaria, et al.
Pubblicazione: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
Empowering Whisper as a Joint Multi-Talker and Target-Talker Speech Recognition System
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
di: Meng, Lingwei, et al.
Pubblicazione: (2024)
SA-SOT: Speaker-Aware Serialized Output Training for Multi-Talker ASR
di: Fan, Zhiyun, et al.
Pubblicazione: (2024)
di: Fan, Zhiyun, et al.
Pubblicazione: (2024)
Semi-Autoregressive Streaming ASR With Label Context
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
di: Arora, Siddhant, et al.
Pubblicazione: (2023)
CALM: Joint Contextual Acoustic-Linguistic Modeling for Personalization of Multi-Speaker ASR
di: Shakeel, Muhammad, et al.
Pubblicazione: (2026)
di: Shakeel, Muhammad, et al.
Pubblicazione: (2026)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
di: Sudo, Yui, et al.
Pubblicazione: (2024)
di: Sudo, Yui, et al.
Pubblicazione: (2024)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
di: Guo, Pengcheng, et al.
Pubblicazione: (2024)
di: Guo, Pengcheng, et al.
Pubblicazione: (2024)
Multi-blank Transducers for Speech Recognition
di: Xu, Hainan, et al.
Pubblicazione: (2022)
di: Xu, Hainan, et al.
Pubblicazione: (2022)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
di: Sudo, Yui, et al.
Pubblicazione: (2025)
di: Sudo, Yui, et al.
Pubblicazione: (2025)
The DKU System for Multi-Speaker Automatic Speech Recognition in MLC-SLM Challenge
di: Lin, Yuke, et al.
Pubblicazione: (2025)
di: Lin, Yuke, et al.
Pubblicazione: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
di: Kocour, Martin, et al.
Pubblicazione: (2025)
di: Kocour, Martin, et al.
Pubblicazione: (2025)
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
di: Someki, Masao, et al.
Pubblicazione: (2023)
di: Someki, Masao, et al.
Pubblicazione: (2023)
Fusion of Discrete Representations and Self-Augmented Representations for Multilingual Automatic Speech Recognition
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
di: Wang, Shih-heng, et al.
Pubblicazione: (2024)
AISHELL-5: The First Open-Source In-Car Multi-Channel Multi-Speaker Speech Dataset for Automatic Speech Diarization and Recognition
di: Dai, Yuhang, et al.
Pubblicazione: (2025)
di: Dai, Yuhang, et al.
Pubblicazione: (2025)
MMM: Multi-Layer Multi-Residual Multi-Stream Discrete Speech Representation from Self-supervised Learning Model
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
di: Arora, Siddhant, et al.
Pubblicazione: (2025)
Chain-Talker: Chain Understanding and Rendering for Empathetic Conversational Speech Synthesis
di: Hu, Yifan, et al.
Pubblicazione: (2025)
di: Hu, Yifan, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
di: Tsunoo, Emiru, et al.
Pubblicazione: (2023) -
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
di: Tsunoo, Emiru, et al.
Pubblicazione: (2025) -
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2024) -
Differentiable K-means for Fully-optimized Discrete Token-based ASR
di: Onda, Kentaro, et al.
Pubblicazione: (2025) -
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
di: Futami, Hayato, et al.
Pubblicazione: (2025)