Transducer-Llama: Integrating LLMs into Streamable Transducer-based Speech Recognition
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Deng, Keqi, Guo, Jinxi, Ma, Yingyi, Moritz, Niko, Woodland, Philip C., Kalinli, Ozlem, Seltzer, Mike |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
par: Deng, Keqi, et autres
Publié: (2023)
par: Deng, Keqi, et autres
Publié: (2023)
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
par: Deng, Keqi, et autres
Publié: (2024)
par: Deng, Keqi, et autres
Publié: (2024)
Effective internal language model training and fusion for factorized transducer model
par: Guo, Jinxi, et autres
Publié: (2024)
par: Guo, Jinxi, et autres
Publié: (2024)
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
par: Ma, Yingyi, et autres
Publié: (2024)
par: Ma, Yingyi, et autres
Publié: (2024)
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
par: Cui, Mingyu, et autres
Publié: (2025)
par: Cui, Mingyu, et autres
Publié: (2025)
SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation
par: Deng, Keqi, et autres
Publié: (2025)
par: Deng, Keqi, et autres
Publié: (2025)
Wav2Prompt: End-to-End Speech Prompt Generation and Tuning For LLM in Zero and Few-shot Learning
par: Deng, Keqi, et autres
Publié: (2024)
par: Deng, Keqi, et autres
Publié: (2024)
Multi-blank Transducers for Speech Recognition
par: Xu, Hainan, et autres
Publié: (2022)
par: Xu, Hainan, et autres
Publié: (2022)
Advanced Long-Content Speech Recognition With Factorized Neural Transducer
par: Gong, Xun, et autres
Publié: (2024)
par: Gong, Xun, et autres
Publié: (2024)
Emotion Neural Transducer for Fine-Grained Speech Emotion Recognition
par: Shen, Siyuan, et autres
Publié: (2024)
par: Shen, Siyuan, et autres
Publié: (2024)
Towards Maximum Likelihood Training for Transducer-based Streaming Speech Recognition
par: Lee, Hyeonseung, et autres
Publié: (2024)
par: Lee, Hyeonseung, et autres
Publié: (2024)
Can Speech LLMs Think while Listening?
par: Shih, Yi-Jen, et autres
Publié: (2025)
par: Shih, Yi-Jen, et autres
Publié: (2025)
CIF-T: A Novel CIF-based Transducer Architecture for Automatic Speech Recognition
par: Zhang, Tian-Hao, et autres
Publié: (2023)
par: Zhang, Tian-Hao, et autres
Publié: (2023)
Transducer Consistency Regularization for Speech to Text Applications
par: Tseng, Cindy, et autres
Publié: (2024)
par: Tseng, Cindy, et autres
Publié: (2024)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
par: Le, Khanh, et autres
Publié: (2025)
par: Le, Khanh, et autres
Publié: (2025)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
par: Bataev, Vladimir, et autres
Publié: (2025)
par: Bataev, Vladimir, et autres
Publié: (2025)
VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
par: Du, Chenpeng, et autres
Publié: (2024)
par: Du, Chenpeng, et autres
Publié: (2024)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
par: Xu, Hainan, et autres
Publié: (2024)
par: Xu, Hainan, et autres
Publié: (2024)
HENT-SRT: Hierarchical Efficient Neural Transducer with Self-Distillation for Joint Speech Recognition and Translation
par: Hussein, Amir, et autres
Publié: (2025)
par: Hussein, Amir, et autres
Publié: (2025)
Faster Speech-LLaMA Inference with Multi-token Prediction
par: Raj, Desh, et autres
Publié: (2024)
par: Raj, Desh, et autres
Publié: (2024)
RNN-Transducer-based Losses for Speech Recognition on Noisy Targets
par: Bataev, Vladimir
Publié: (2025)
par: Bataev, Vladimir
Publié: (2025)
XLSR-Transducer: Streaming ASR for Self-Supervised Pretrained Models
par: Kumar, Shashi, et autres
Publié: (2024)
par: Kumar, Shashi, et autres
Publié: (2024)
M-BEST-RQ: A Multi-Channel Speech Foundation Model for Smart Glasses
par: Yang, Yufeng, et autres
Publié: (2024)
par: Yang, Yufeng, et autres
Publié: (2024)
High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model
par: Lee, Joun Yeop, et autres
Publié: (2024)
par: Lee, Joun Yeop, et autres
Publié: (2024)
CUSIDE-T: Chunking, Simulating Future and Decoding for Transducer based Streaming ASR
par: Zhao, Wenbo, et autres
Publié: (2024)
par: Zhao, Wenbo, et autres
Publié: (2024)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
par: Zhou, Wei, et autres
Publié: (2024)
par: Zhou, Wei, et autres
Publié: (2024)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
par: Kumar, Shashi, et autres
Publié: (2024)
par: Kumar, Shashi, et autres
Publié: (2024)
Diffusion-based Generative Modeling with Discriminative Guidance for Streamable Speech Enhancement
par: Li, Chenda, et autres
Publié: (2024)
par: Li, Chenda, et autres
Publié: (2024)
Masked Self-distilled Transducer-based Keyword Spotting with Semi-autoregressive Decoding
par: Xi, Yu, et autres
Publié: (2025)
par: Xi, Yu, et autres
Publié: (2025)
Incorporating Class-based Language Model for Named Entity Recognition in Factorized Neural Transducer
par: Wang, Peng, et autres
Publié: (2023)
par: Wang, Peng, et autres
Publié: (2023)
Conversational Speech Naturalness Predictor
par: Xu, Anfeng, et autres
Publié: (2026)
par: Xu, Anfeng, et autres
Publié: (2026)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
par: Sudo, Yui, et autres
Publié: (2024)
par: Sudo, Yui, et autres
Publié: (2024)
Promptformer: Prompted Conformer Transducer for ASR
par: Duarte-Torres, Sergio, et autres
Publié: (2024)
par: Duarte-Torres, Sergio, et autres
Publié: (2024)
FireRedTTS-1S: An Upgraded Streamable Foundation Text-to-Speech System
par: Guo, Hao-Han, et autres
Publié: (2025)
par: Guo, Hao-Han, et autres
Publié: (2025)
A High-Quality and Low-Complexity Streamable Neural Speech Codec with Knowledge Distillation
par: Zhang, En-Wei, et autres
Publié: (2025)
par: Zhang, En-Wei, et autres
Publié: (2025)
Alignment-Free Training for Transducer-based Multi-Talker ASR
par: Moriya, Takafumi, et autres
Publié: (2024)
par: Moriya, Takafumi, et autres
Publié: (2024)
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
par: Lashkarashvili, Nineli, et autres
Publié: (2024)
par: Lashkarashvili, Nineli, et autres
Publié: (2024)
Distribution-based Emotion Recognition in Conversation
par: Wu, Wen, et autres
Publié: (2022)
par: Wu, Wen, et autres
Publié: (2022)
TDT-KWS: Fast And Accurate Keyword Spotting Using Token-and-duration Transducer
par: Xi, Yu, et autres
Publié: (2024)
par: Xi, Yu, et autres
Publié: (2024)
Lightweight Transducer Based on Frame-Level Criterion
par: Wan, Genshun, et autres
Publié: (2024)
par: Wan, Genshun, et autres
Publié: (2024)
Documents similaires
-
Label-Synchronous Neural Transducer for Adaptable Online E2E Speech Recognition
par: Deng, Keqi, et autres
Publié: (2023) -
Label-Synchronous Neural Transducer for E2E Simultaneous Speech Translation
par: Deng, Keqi, et autres
Publié: (2024) -
Effective internal language model training and fusion for factorized transducer model
par: Guo, Jinxi, et autres
Publié: (2024) -
Effective Text Adaptation for LLM-based ASR through Soft Prompt Fine-Tuning
par: Ma, Yingyi, et autres
Publié: (2024) -
Exploring Cross-Utterance Speech Contexts for Conformer-Transducer Speech Recognition Systems
par: Cui, Mingyu, et autres
Publié: (2025)