CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Hanwen, Yusuyin, Saierdaer, Huang, Hao, Ou, Zhijian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
von: Dong, Lukuang, et al.
Veröffentlicht: (2026)
Enhancing CTC-based speech recognition with diverse modeling units
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2025)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2025)
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)
CTC-aligned Audio-Text Embedding for Streaming Open-vocabulary Keyword Spotting
von: Jin, Sichen, et al.
Veröffentlicht: (2024)
von: Jin, Sichen, et al.
Veröffentlicht: (2024)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
von: Kang, Jiawen, et al.
Veröffentlicht: (2024)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
TTS-CtrlNet: Time varying emotion aligned text-to-speech generation with ControlNet
von: Jeong, Jaeseok, et al.
Veröffentlicht: (2025)
von: Jeong, Jaeseok, et al.
Veröffentlicht: (2025)
Phonemes vs. Projectors: An Investigation of Speech-Language Interfaces for LLM-based ASR
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
von: Li, Ziwei, et al.
Veröffentlicht: (2026)
Phoneme-based speech recognition driven by large language models and sampling marginalization
von: Ma, Te, et al.
Veröffentlicht: (2025)
von: Ma, Te, et al.
Veröffentlicht: (2025)
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
Differentiable Reward Optimization for LLM based TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
Explore the Reinforcement Learning for the LLM based ASR and TTS system
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
von: Gao, Changfeng, et al.
Veröffentlicht: (2025)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
RRPO: Robust Reward Policy Optimization for LLM-based Emotional TTS
von: Wang, Cong, et al.
Veröffentlicht: (2025)
von: Wang, Cong, et al.
Veröffentlicht: (2025)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
Unimodal Aggregation for CTC-based Speech Recognition
von: Fang, Ying, et al.
Veröffentlicht: (2023)
von: Fang, Ying, et al.
Veröffentlicht: (2023)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
Towards Lightweight and Stable Zero-shot TTS with Self-distilled Representation Disentanglement
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
von: Chen, Qianniu, et al.
Veröffentlicht: (2025)
LegoSLM: Connecting LLM with Speech Encoder using CTC Posteriors
von: Ma, Rao, et al.
Veröffentlicht: (2025)
von: Ma, Rao, et al.
Veröffentlicht: (2025)
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information
von: Wang, Rui, et al.
Veröffentlicht: (2025)
von: Wang, Rui, et al.
Veröffentlicht: (2025)
Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
von: Wang, Xinsheng, et al.
Veröffentlicht: (2025)
SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS
von: Singh, Ayush Pratap, et al.
Veröffentlicht: (2026)
von: Singh, Ayush Pratap, et al.
Veröffentlicht: (2026)
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
von: Altinok, Duygu
Veröffentlicht: (2025)
von: Altinok, Duygu
Veröffentlicht: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2025)
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2025)
Multilingual Audio-Visual Speech Recognition with Hybrid CTC/RNN-T Fast Conformer
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
von: Burchi, Maxime, et al.
Veröffentlicht: (2024)
On the Effectiveness of Acoustic BPE in Decoder-Only TTS
von: Li, Bohan, et al.
Veröffentlicht: (2024)
von: Li, Bohan, et al.
Veröffentlicht: (2024)
ReStyle-TTS: Relative and Continuous Style Control for Zero-Shot Speech Synthesis
von: Li, Haitao, et al.
Veröffentlicht: (2026)
von: Li, Haitao, et al.
Veröffentlicht: (2026)
A Non-autoregressive Model for Joint STT and TTS
von: Sunder, Vishal, et al.
Veröffentlicht: (2025)
von: Sunder, Vishal, et al.
Veröffentlicht: (2025)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Unifying Streaming and Non-streaming Zipformer-based ASR
von: Sharma, Bidisha, et al.
Veröffentlicht: (2025)
von: Sharma, Bidisha, et al.
Veröffentlicht: (2025)
CTC-Assisted LLM-Based Contextual ASR
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM-based phoneme-to-grapheme for phoneme-based speech recognition
von: Ma, Te, et al.
Veröffentlicht: (2025) -
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
von: Dong, Lukuang, et al.
Veröffentlicht: (2026) -
Enhancing CTC-based speech recognition with diverse modeling units
von: Han, Shiyi, et al.
Veröffentlicht: (2024) -
Pronunciation-Lexicon Free Training for Phoneme-based Crosslingual ASR via Joint Stochastic Approximation
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2025) -
Whistle: Data-Efficient Multilingual and Crosslingual Speech Recognition via Weakly Phonetic Supervision
von: Yusuyin, Saierdaer, et al.
Veröffentlicht: (2024)