Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Hono, Yukiya, Mitsuda, Koh, Zhao, Tianyu, Mitsui, Kentaro, Wakatsuki, Toshiaki, Sawada, Kei |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
par: Mitsui, Kentaro, et autres
Publié: (2024)
par: Mitsui, Kentaro, et autres
Publié: (2024)
Release of Pre-Trained Models for the Japanese Language
par: Sawada, Kei, et autres
Publié: (2024)
par: Sawada, Kei, et autres
Publié: (2024)
End-to-End Speech Recognition with Pre-trained Masked Language Model
par: Higuchi, Yosuke, et autres
Publié: (2024)
par: Higuchi, Yosuke, et autres
Publié: (2024)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
par: Rufai, Amina Mardiyyah, et autres
Publié: (2020)
par: Rufai, Amina Mardiyyah, et autres
Publié: (2020)
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
par: Wang, Chen, et autres
Publié: (2025)
par: Wang, Chen, et autres
Publié: (2025)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
par: He, Xinlu, et autres
Publié: (2025)
par: He, Xinlu, et autres
Publié: (2025)
An End-to-End Speech Summarization Using Large Language Model
par: Shang, Hengchao, et autres
Publié: (2024)
par: Shang, Hengchao, et autres
Publié: (2024)
A Non-autoregressive Generation Framework for End-to-End Simultaneous Speech-to-Speech Translation
par: Ma, Zhengrui, et autres
Publié: (2024)
par: Ma, Zhengrui, et autres
Publié: (2024)
Recent Advances in End-to-End Simultaneous Speech Translation
par: Liu, Xiaoqian, et autres
Publié: (2024)
par: Liu, Xiaoqian, et autres
Publié: (2024)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
par: Higuchi, Yosuke, et autres
Publié: (2023)
par: Higuchi, Yosuke, et autres
Publié: (2023)
Representation Purification for End-to-End Speech Translation
par: Zhang, Chengwei, et autres
Publié: (2024)
par: Zhang, Chengwei, et autres
Publié: (2024)
End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach
par: Abdullah, Abdulhady Abas, et autres
Publié: (2024)
par: Abdullah, Abdulhady Abas, et autres
Publié: (2024)
TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer
par: Bataev, Vladimir, et autres
Publié: (2025)
par: Bataev, Vladimir, et autres
Publié: (2025)
Data Augmentation for End-to-end Code-switching Speech Recognition
par: Du, Chenpeng, et autres
Publié: (2020)
par: Du, Chenpeng, et autres
Publié: (2020)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
par: Tsunoo, Emiru, et autres
Publié: (2024)
par: Tsunoo, Emiru, et autres
Publié: (2024)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
par: Farhadipour, Aref, et autres
Publié: (2023)
par: Farhadipour, Aref, et autres
Publié: (2023)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
par: Huang, Wuwei, et autres
Publié: (2025)
par: Huang, Wuwei, et autres
Publié: (2025)
StutterZero and StutterFormer: End-to-End Speech Conversion for Stuttering Transcription and Correction
par: Xu, Qianheng
Publié: (2025)
par: Xu, Qianheng
Publié: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
par: Kocour, Martin, et autres
Publié: (2025)
par: Kocour, Martin, et autres
Publié: (2025)
Continual Learning for Monolingual End-to-End Automatic Speech Recognition
par: Eeckt, Steven Vander, et autres
Publié: (2021)
par: Eeckt, Steven Vander, et autres
Publié: (2021)
When End-to-End is Overkill: Rethinking Cascaded Speech-to-Text Translation
par: Min, Anna, et autres
Publié: (2025)
par: Min, Anna, et autres
Publié: (2025)
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
par: Pothula, Aishwarya, et autres
Publié: (2025)
par: Pothula, Aishwarya, et autres
Publié: (2025)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
par: Agro, Maha Tufail, et autres
Publié: (2025)
par: Agro, Maha Tufail, et autres
Publié: (2025)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
par: Liu, Heyang, et autres
Publié: (2024)
par: Liu, Heyang, et autres
Publié: (2024)
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
par: Polák, Peter, et autres
Publié: (2023)
par: Polák, Peter, et autres
Publié: (2023)
Disentangled-Transformer: An Explainable End-to-End Automatic Speech Recognition Model with Speech Content-Context Separation
par: Wang, Pu, et autres
Publié: (2024)
par: Wang, Pu, et autres
Publié: (2024)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
par: Subramanian, Aswin Shanmugam, et autres
Publié: (2025)
par: Subramanian, Aswin Shanmugam, et autres
Publié: (2025)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
par: Hirano, Yuta, et autres
Publié: (2025)
par: Hirano, Yuta, et autres
Publié: (2025)
PeriodGrad: Towards Pitch-Controllable Neural Vocoder Based on a Diffusion Probabilistic Model
par: Hono, Yukiya, et autres
Publié: (2024)
par: Hono, Yukiya, et autres
Publié: (2024)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
par: Wang, Peidong, et autres
Publié: (2024)
par: Wang, Peidong, et autres
Publié: (2024)
End-to-End Speech-to-Text Translation: A Survey
par: Sethiya, Nivedita, et autres
Publié: (2023)
par: Sethiya, Nivedita, et autres
Publié: (2023)
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
par: Zhou, Xuanru, et autres
Publié: (2024)
par: Zhou, Xuanru, et autres
Publié: (2024)
Sentence-wise Speech Summarization: Task, Datasets, and End-to-End Modeling with LM Knowledge Distillation
par: Matsuura, Kohei, et autres
Publié: (2024)
par: Matsuura, Kohei, et autres
Publié: (2024)
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
par: Hsu, Ming-Hao, et autres
Publié: (2026)
par: Hsu, Ming-Hao, et autres
Publié: (2026)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
par: Vendrame, Katia, et autres
Publié: (2025)
par: Vendrame, Katia, et autres
Publié: (2025)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
par: Shakeel, Muhammad, et autres
Publié: (2024)
par: Shakeel, Muhammad, et autres
Publié: (2024)
Privacy-Preserving End-to-End Full-Duplex Speech Dialogue Models
par: Kuzmin, Nikita, et autres
Publié: (2026)
par: Kuzmin, Nikita, et autres
Publié: (2026)
Qifusion-Net: Layer-adapted Stream/Non-stream Model for End-to-End Multi-Accent Speech Recognition
par: Chen, Jinming, et autres
Publié: (2024)
par: Chen, Jinming, et autres
Publié: (2024)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
par: Li, Tianpeng, et autres
Publié: (2025)
par: Li, Tianpeng, et autres
Publié: (2025)
Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
par: Eeckt, Steven Vander, et autres
Publié: (2022)
par: Eeckt, Steven Vander, et autres
Publié: (2022)
Documents similaires
-
PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems
par: Mitsui, Kentaro, et autres
Publié: (2024) -
Release of Pre-Trained Models for the Japanese Language
par: Sawada, Kei, et autres
Publié: (2024) -
End-to-End Speech Recognition with Pre-trained Masked Language Model
par: Higuchi, Yosuke, et autres
Publié: (2024) -
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
par: Rufai, Amina Mardiyyah, et autres
Publié: (2020) -
OpenS2S: Advancing Fully Open-Source End-to-End Empathetic Large Speech Language Model
par: Wang, Chen, et autres
Publié: (2025)