LV-CTC: Non-autoregressive ASR with CTC and latent variable models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fujita, Yuya, Watanabe, Shinji, Chang, Xuankai, Maekaku, Takashi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023)
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
von: Saon, George, et al.
Veröffentlicht: (2026)
von: Saon, George, et al.
Veröffentlicht: (2026)
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026)
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)
SQ-Whisper: Speaker-Querying based Whisper Model for Target-Speaker ASR
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
von: Guo, Pengcheng, et al.
Veröffentlicht: (2024)
CTC-Assisted LLM-Based Contextual ASR
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
von: Yang, Guanrou, et al.
Veröffentlicht: (2024)
Fast Context-Biasing for CTC and Transducer ASR models with CTC-based Word Spotter
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
von: Andrusenko, Andrei, et al.
Veröffentlicht: (2024)
CTC Blank Triggered Dynamic Layer-Skipping for Efficient CTC-based Speech Recognition
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
von: Hou, Junfeng, et al.
Veröffentlicht: (2024)
Joint Beam Search Integrating CTC, Attention, and Transducer Decoders
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
von: Sudo, Yui, et al.
Veröffentlicht: (2024)
Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR
von: Li, Longhao, et al.
Veröffentlicht: (2025)
von: Li, Longhao, et al.
Veröffentlicht: (2025)
NLE: Non-autoregressive LLM-based ASR by Transcript Editing
von: Dekel, Avihu, et al.
Veröffentlicht: (2026)
von: Dekel, Avihu, et al.
Veröffentlicht: (2026)
CR-CTC: Consistency regularization on CTC for improved speech recognition
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
von: Yao, Zengwei, et al.
Veröffentlicht: (2024)
CTC-based Non-autoregressive Textless Speech-to-Speech Translation
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
von: Fang, Qingkai, et al.
Veröffentlicht: (2024)
SiamCTC: Learning Speech Representations through Monotonic Temporal Alignment
von: Eom, SooHwan, et al.
Veröffentlicht: (2026)
von: Eom, SooHwan, et al.
Veröffentlicht: (2026)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Evaluating Self-Supervised Speech Models via Text-Based LLMS
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
von: Maekaku, Takashi, et al.
Veröffentlicht: (2025)
CTC-TTS: LLM-based dual-streaming text-to-speech with CTC alignment
von: Liu, Hanwen, et al.
Veröffentlicht: (2026)
von: Liu, Hanwen, et al.
Veröffentlicht: (2026)
Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization
von: Altinok, Duygu
Veröffentlicht: (2025)
von: Altinok, Duygu
Veröffentlicht: (2025)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
CJST: CTC Compressor based Joint Speech and Text Training for Decoder-Only ASR
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
von: Zhou, Wei, et al.
Veröffentlicht: (2024)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
von: Sakuma, Asahi, et al.
Veröffentlicht: (2025)
LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
von: Ye, Zongli, et al.
Veröffentlicht: (2025)
CTC-GMM: CTC guided modality matching for fast and accurate streaming speech translation
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
von: Zhao, Rui, et al.
Veröffentlicht: (2024)
NTC-KWS: Noise-aware CTC for Robust Keyword Spotting
von: Xi, Yu, et al.
Veröffentlicht: (2024)
von: Xi, Yu, et al.
Veröffentlicht: (2024)
AdaMER-CTC: Connectionist Temporal Classification with Adaptive Maximum Entropy Regularization for Automatic Speech Recognition
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
von: Eom, SooHwan, et al.
Veröffentlicht: (2024)
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
von: Zheng, Yuang, et al.
Veröffentlicht: (2026)
von: Zheng, Yuang, et al.
Veröffentlicht: (2026)
Enhancing Code-Switching ASR Leveraging Non-Peaky CTC Loss and Deep Language Posterior Injection
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
von: Yang, Tzu-Ting, et al.
Veröffentlicht: (2024)
Enhancing CTC-based speech recognition with diverse modeling units
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
von: Han, Shiyi, et al.
Veröffentlicht: (2024)
Unimodal Aggregation for CTC-based Speech Recognition
von: Fang, Ying, et al.
Veröffentlicht: (2023)
von: Fang, Ying, et al.
Veröffentlicht: (2023)
Improving Zero-Shot Chinese-English Code-Switching ASR with kNN-CTC and Gated Monolingual Datastores
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
von: Zhou, Jiaming, et al.
Veröffentlicht: (2024)
Align-Consistency: Improving Non-autoregressive and Semi-supervised ASR with Consistency Regularization
von: Huang, Wanting, et al.
Veröffentlicht: (2026)
von: Huang, Wanting, et al.
Veröffentlicht: (2026)
Automatic Speech Recognition with BERT and CTC Transformers: A Review
von: Djeffal, Noussaiba, et al.
Veröffentlicht: (2024)
von: Djeffal, Noussaiba, et al.
Veröffentlicht: (2024)
Enhancing GOP in CTC-Based Mispronunciation Detection with Phonological Knowledge
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025)
von: Parikh, Aditya Kamlesh, et al.
Veröffentlicht: (2025)
CC-G2PnP: Streaming Grapheme-to-Phoneme and prosody with Conformer-CTC for unsegmented languages
von: Shirahata, Yuma, et al.
Veröffentlicht: (2026)
von: Shirahata, Yuma, et al.
Veröffentlicht: (2026)
Analyzing the Importance of Blank for CTC-Based Knowledge Distillation
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2025)
von: Hilmes, Benedikt, et al.
Veröffentlicht: (2025)
WCTC-Biasing: Retraining-free Contextual Biasing ASR with Wildcard CTC-based Keyword Spotting and Inter-layer Biasing
von: Nakagome, Yu, et al.
Veröffentlicht: (2025)
von: Nakagome, Yu, et al.
Veröffentlicht: (2025)
Less Peaky and More Accurate CTC Forced Alignment by Label Priors
von: Huang, Ruizhe, et al.
Veröffentlicht: (2024)
von: Huang, Ruizhe, et al.
Veröffentlicht: (2024)
FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
von: Grigoryan, Lilit, et al.
Veröffentlicht: (2025)
SegAug: CTC-Aligned Segmented Augmentation For Robust RNN-Transducer Based Speech Recognition
von: Le, Khanh, et al.
Veröffentlicht: (2025)
von: Le, Khanh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels
von: Zhou, Jiaming, et al.
Veröffentlicht: (2023) -
Self-Speculative Decoding for LLM-based ASR with CTC Encoder Drafts
von: Saon, George, et al.
Veröffentlicht: (2026) -
Contextual Biasing for Streaming ASR via CTC-based Word Spotting
von: Tsai, Kai-Chen, et al.
Veröffentlicht: (2026) -
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2023) -
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
von: Moriya, Takafumi, et al.
Veröffentlicht: (2025)