Gespeichert in:
| Hauptverfasser: | Du, Chenpeng, Li, Hao, Lu, Yizhou, Wang, Lan, Qian, Yanmin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2020
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2011.02160 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Decoder-only Architecture for Streaming End-to-end Speech Recognition
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023)
von: Shao, Hang, et al.
Veröffentlicht: (2023)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition
von: Xia, Yinfeng, et al.
Veröffentlicht: (2025)
von: Xia, Yinfeng, et al.
Veröffentlicht: (2025)
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
von: Wang, Hankun, et al.
Veröffentlicht: (2024)
Adapting Whisper for Parameter-efficient Code-Switching Speech Recognition via Soft Prompt Tuning
von: Yang, Hongli, et al.
Veröffentlicht: (2025)
von: Yang, Hongli, et al.
Veröffentlicht: (2025)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
Measuring Entrainment in Spontaneous Code-switched Speech
von: Bhattacharya, Debasmita, et al.
Veröffentlicht: (2023)
von: Bhattacharya, Debasmita, et al.
Veröffentlicht: (2023)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
von: Shen, Peng, et al.
Veröffentlicht: (2025)
von: Shen, Peng, et al.
Veröffentlicht: (2025)
RAG-Boost: Retrieval-Augmented Generation Enhanced LLM-based Speech Recognition
von: Wang, Pengcheng, et al.
Veröffentlicht: (2025)
von: Wang, Pengcheng, et al.
Veröffentlicht: (2025)
SLIDE: Integrating Speech Language Model with LLM for Spontaneous Spoken Dialogue Generation
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
von: Lu, Haitian, et al.
Veröffentlicht: (2025)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
SimulMEGA: MoE Routers are Advanced Policy Makers for Simultaneous Speech Translation
von: Le, Chenyang, et al.
Veröffentlicht: (2025)
von: Le, Chenyang, et al.
Veröffentlicht: (2025)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2024)
Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
von: Žavoronkov, Aleksei, et al.
Veröffentlicht: (2025)
von: Žavoronkov, Aleksei, et al.
Veröffentlicht: (2025)
SpeechColab Leaderboard: An Open-Source Platform for Automatic Speech Recognition Evaluation
von: Du, Jiayu, et al.
Veröffentlicht: (2024)
von: Du, Jiayu, et al.
Veröffentlicht: (2024)
Novel Parasitic Dual-Scale Modeling for Efficient and Accurate Multilingual Speech Translation
von: Le, Chenyang, et al.
Veröffentlicht: (2025)
von: Le, Chenyang, et al.
Veröffentlicht: (2025)
Improving Accented Speech Recognition using Data Augmentation based on Unsupervised Text-to-Speech Synthesis
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
von: Do, Cong-Thanh, et al.
Veröffentlicht: (2024)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
Anatomy of the Modality Gap: Dissecting the Internal States of End-to-End Speech LLMs
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2026)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2026)
Gammatonegram Representation for End-to-End Dysarthric Speech Processing Tasks: Speech Recognition, Speaker Identification, and Intelligibility Assessment
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
von: Farhadipour, Aref, et al.
Veröffentlicht: (2023)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
YOLO-Stutter: End-to-end Region-Wise Speech Dysfluency Detection
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
von: Zhou, Xuanru, et al.
Veröffentlicht: (2024)
End-to-End Speech Translation for Low-Resource Languages Using Weakly Labeled Data
von: Pothula, Aishwarya, et al.
Veröffentlicht: (2025)
von: Pothula, Aishwarya, et al.
Veröffentlicht: (2025)
Continual Learning for Monolingual End-to-End Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2021)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2021)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
Enhancing Code-switched Text-to-Speech Synthesis Capability in Large Language Models with only Monolingual Corpora
von: Xu, Jing, et al.
Veröffentlicht: (2024)
von: Xu, Jing, et al.
Veröffentlicht: (2024)
Dynamic Data Pruning for Automatic Speech Recognition
von: Xiao, Qiao, et al.
Veröffentlicht: (2024)
von: Xiao, Qiao, et al.
Veröffentlicht: (2024)
Optimizing Byte-level Representation for End-to-end ASR
von: Hsiao, Roger, et al.
Veröffentlicht: (2024)
von: Hsiao, Roger, et al.
Veröffentlicht: (2024)
Learning Speech Representations with Variational Predictive Coding
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2025)
von: Yeh, Sung-Lin, et al.
Veröffentlicht: (2025)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
Dolphin: A Large-Scale Automatic Speech Recognition Model for Eastern Languages
von: Meng, Yangyang, et al.
Veröffentlicht: (2025)
von: Meng, Yangyang, et al.
Veröffentlicht: (2025)
Enhancing Code-Switching Speech Recognition with LID-Based Collaborative Mixture of Experts Model
von: Huang, Hukai, et al.
Veröffentlicht: (2024)
von: Huang, Hukai, et al.
Veröffentlicht: (2024)
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers
von: Hentschel, Michael, et al.
Veröffentlicht: (2024)
von: Hentschel, Michael, et al.
Veröffentlicht: (2024)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
von: Moslem, Yasmin
Veröffentlicht: (2024)
von: Moslem, Yasmin
Veröffentlicht: (2024)
Predictive Speech Recognition and End-of-Utterance Detection Towards Spoken Dialog Systems
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
von: Zink, Oswald, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Decoder-only Architecture for Streaming End-to-end Speech Recognition
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024) -
DQ-Whisper: Joint Distillation and Quantization for Efficient Multilingual Speech Recognition
von: Shao, Hang, et al.
Veröffentlicht: (2023) -
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024) -
MFLA: Monotonic Finite Look-ahead Attention for Streaming Speech Recognition
von: Xia, Yinfeng, et al.
Veröffentlicht: (2025) -
Why Do Speech Language Models Fail to Generate Semantically Coherent Outputs? A Modality Evolving Perspective
von: Wang, Hankun, et al.
Veröffentlicht: (2024)