End-to-End Simultaneous Dysarthric Speech Reconstruction with Frame-Level Adaptor and Multiple Wait-k Knowledge Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Minghui, Tang, Haitao, Fan, Jiahuan, Liao, Ruizhi, Zhang, Yanyong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DARS: Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement
von: Wu, Minghui, et al.
Veröffentlicht: (2026)
von: Wu, Minghui, et al.
Veröffentlicht: (2026)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Recent Advances in End-to-End Simultaneous Speech Translation
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024)
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024)
UNIT-DSR: Dysarthric Speech Reconstruction System Using Speech Unit Normalization
von: Wang, Yuejiao, et al.
Veröffentlicht: (2024)
von: Wang, Yuejiao, et al.
Veröffentlicht: (2024)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
Prototype-Based Disentanglement for Controllable Dysarthric Speech Synthesis
von: Wang, Haoshen, et al.
Veröffentlicht: (2026)
von: Wang, Haoshen, et al.
Veröffentlicht: (2026)
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers
von: Hentschel, Michael, et al.
Veröffentlicht: (2024)
von: Hentschel, Michael, et al.
Veröffentlicht: (2024)
End-to-End Speech-to-Text Translation: A Survey
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
A Few-Shot Approach to Dysarthric Speech Intelligibility Level Classification Using Transformers
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023)
von: Chowdary, Paleti Nikhil, et al.
Veröffentlicht: (2023)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
von: Moslem, Yasmin
Veröffentlicht: (2024)
von: Moslem, Yasmin
Veröffentlicht: (2024)
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
BLSP-KD: Bootstrapping Language-Speech Pre-training via Knowledge Distillation
von: Wang, Chen, et al.
Veröffentlicht: (2024)
von: Wang, Chen, et al.
Veröffentlicht: (2024)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
von: Bataev, Vladimir, et al.
Veröffentlicht: (2025)
Streaming Bilingual End-to-End ASR model using Attention over Multiple Softmax
von: Patil, Aditya, et al.
Veröffentlicht: (2024)
von: Patil, Aditya, et al.
Veröffentlicht: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
Applications of Artificial Intelligence for Cross-language Intelligibility Assessment of Dysarthric Speech
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
von: Yeo, Eunjung, et al.
Veröffentlicht: (2025)
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
A Calculus-Based Framework for Determining Vocabulary Size in End-to-End ASR
von: Kopparapu, Sunil Kumar
Veröffentlicht: (2026)
von: Kopparapu, Sunil Kumar
Veröffentlicht: (2026)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
von: Wang, Chen, et al.
Veröffentlicht: (2025)
von: Wang, Chen, et al.
Veröffentlicht: (2025)
End-to-end Contrastive Language-Speech Pretraining Model For Long-form Spoken Question Answering
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
SpeechDPR: End-to-End Spoken Passage Retrieval for Open-Domain Spoken Question Answering
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
von: Lin, Chyi-Jiunn, et al.
Veröffentlicht: (2024)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
von: Polák, Peter, et al.
Veröffentlicht: (2023)
von: Polák, Peter, et al.
Veröffentlicht: (2023)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
von: Lee, Wonjun, et al.
Veröffentlicht: (2025)
Idiosyncratic Versus Normative Modeling of Atypical Speech Recognition: Dysarthric Case Studies
von: Raja, Vishnu, et al.
Veröffentlicht: (2025)
von: Raja, Vishnu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DARS: Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement
von: Wu, Minghui, et al.
Veröffentlicht: (2026) -
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023) -
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025) -
Recent Advances in End-to-End Simultaneous Speech Translation
von: Liu, Xiaoqian, et al.
Veröffentlicht: (2024) -
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
von: Ma, Zhengrui, et al.
Veröffentlicht: (2024)