Streaming Bilingual End-to-End ASR model using Attention over Multiple Softmax
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Patil, Aditya, Joshi, Vikas, Agrawal, Purvi, Mehta, Rupesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Building English ASR model with regional language support
von: Agrawal, Purvi, et al.
Veröffentlicht: (2025)
von: Agrawal, Purvi, et al.
Veröffentlicht: (2025)
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
von: Li, Chia-Yu, et al.
Veröffentlicht: (2024)
von: Li, Chia-Yu, et al.
Veröffentlicht: (2024)
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
von: Kopparapu, Sunil Kumar, et al.
Veröffentlicht: (2024)
von: Kopparapu, Sunil Kumar, et al.
Veröffentlicht: (2024)
Alternating Weak Triphone/BPE Alignment Supervision from Hybrid Model Improves End-to-End ASR
von: Jiang, Jintao, et al.
Veröffentlicht: (2024)
von: Jiang, Jintao, et al.
Veröffentlicht: (2024)
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data
von: Cui, Can, et al.
Veröffentlicht: (2023)
von: Cui, Can, et al.
Veröffentlicht: (2023)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
Mamba for Streaming ASR Combined with Unimodal Aggregation
von: Fang, Ying, et al.
Veröffentlicht: (2024)
von: Fang, Ying, et al.
Veröffentlicht: (2024)
Semi-Autoregressive Streaming ASR With Label Context
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
von: Arora, Siddhant, et al.
Veröffentlicht: (2023)
End-to-End Speech-to-Text Translation: A Survey
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
von: Sethiya, Nivedita, et al.
Veröffentlicht: (2023)
Speaker Adaptation for Quantised End-to-End ASR Models
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2024)
An End-to-End Speech Summarization Using Large Language Model
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
von: Shang, Hengchao, et al.
Veröffentlicht: (2024)
An investigation of phrase break prediction in an End-to-End TTS system
von: Vadapalli, Anandaswarup
Veröffentlicht: (2023)
von: Vadapalli, Anandaswarup
Veröffentlicht: (2023)
Towards Building an End-to-End Multilingual Automatic Lyrics Transcription Model
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
von: Huang, Jiawen, et al.
Veröffentlicht: (2024)
Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
von: Li, Tianpeng, et al.
Veröffentlicht: (2025)
Scaling and Prompting for Improved End-to-End Spoken Grammatical Error Correction
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
von: Qian, Mengjie, et al.
Veröffentlicht: (2025)
Joint Training And Decoding for Multilingual End-to-End Simultaneous Speech Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
PRoDeliberation: Parallel Robust Deliberation for End-to-End Spoken Language Understanding
von: Le, Trang, et al.
Veröffentlicht: (2024)
von: Le, Trang, et al.
Veröffentlicht: (2024)
Leveraging Synthetic Audio Data for End-to-End Low-Resource Speech Translation
von: Moslem, Yasmin
Veröffentlicht: (2024)
von: Moslem, Yasmin
Veröffentlicht: (2024)
Acoustically Precise Hesitation Tagging Is Essential for End-to-End Verbatim Transcription Systems
von: Lin, Jhen-Ke, et al.
Veröffentlicht: (2025)
von: Lin, Jhen-Ke, et al.
Veröffentlicht: (2025)
Adapting Diarization-Conditioned Whisper for End-to-End Multi-Talker Speech Recognition
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
von: Kocour, Martin, et al.
Veröffentlicht: (2025)
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
von: Zeng, Aohan, et al.
Veröffentlicht: (2024)
Soft Language Identification for Language-Agnostic Many-to-One End-to-End Speech Translation
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
von: Wang, Peidong, et al.
Veröffentlicht: (2024)
Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
von: Liu, Heyang, et al.
Veröffentlicht: (2024)
Beyond Binary: Multiclass Paraphasia Detection with Generative Pretrained Transformers and End-to-End Models
von: Perez, Matthew, et al.
Veröffentlicht: (2024)
von: Perez, Matthew, et al.
Veröffentlicht: (2024)
Code-Switching in End-to-End Automatic Speech Recognition: A Systematic Literature Review
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
von: Agro, Maha Tufail, et al.
Veröffentlicht: (2025)
VoxEval: Benchmarking the Knowledge Understanding Capabilities of End-to-End Spoken Language Models
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
von: Cui, Wenqian, et al.
Veröffentlicht: (2025)
Fast Streaming Transducer ASR Prototyping via Knowledge Distillation with Whisper
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
von: Thorbecke, Iuliia, et al.
Veröffentlicht: (2024)
Speaker-Smoothed kNN Speaker Adaptation for End-to-End ASR
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
von: Li, Shaojun, et al.
Veröffentlicht: (2024)
AdaST: Dynamically Adapting Encoder States in the Decoder for End-to-End Speech-to-Text Translation
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
von: Huang, Wuwei, et al.
Veröffentlicht: (2025)
Joint Speech and Text Training for LLM-Based End-to-End Spoken Dialogue State Tracking
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
von: Vendrame, Katia, et al.
Veröffentlicht: (2025)
Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
von: Huang, Ailin, et al.
Veröffentlicht: (2025)
Improving Practical Aspects of End-to-End Multi-Talker Speech Recognition for Online and Offline Scenarios
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
von: Subramanian, Aswin Shanmugam, et al.
Veröffentlicht: (2025)
SC-SOT: Conditioning the Decoder on Diarized Speaker Information for End-to-End Overlapped Speech Recognition
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
von: Hirano, Yuta, et al.
Veröffentlicht: (2025)
The Sound of Healthcare: Improving Medical Transcription ASR Accuracy with Large Language Models
von: Adedeji, Ayo, et al.
Veröffentlicht: (2024)
von: Adedeji, Ayo, et al.
Veröffentlicht: (2024)
Song Data Cleansing for End-to-End Neural Singer Diarization Using Neural Analysis and Synthesis Framework
von: Munakata, Hokuto, et al.
Veröffentlicht: (2024)
von: Munakata, Hokuto, et al.
Veröffentlicht: (2024)
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
von: Prabhavalkar, Rohit, et al.
Veröffentlicht: (2024)
SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
von: Wang, Peidong, et al.
Veröffentlicht: (2026)
von: Wang, Peidong, et al.
Veröffentlicht: (2026)
Harnessing the Zero-Shot Power of Instruction-Tuned Large Language Model in End-to-End Speech Recognition
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2023)
Enhanced ASR Robustness to Packet Loss with a Front-End Adaptation Network
von: Dissen, Yehoshua, et al.
Veröffentlicht: (2024)
von: Dissen, Yehoshua, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Building English ASR model with regional language support
von: Agrawal, Purvi, et al.
Veröffentlicht: (2025) -
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
von: Li, Chia-Yu, et al.
Veröffentlicht: (2024) -
A cost minimization approach to fix the vocabulary size in a tokenizer for an End-to-End ASR system
von: Kopparapu, Sunil Kumar, et al.
Veröffentlicht: (2024) -
Alternating Weak Triphone/BPE Alignment Supervision from Hybrid Model Improves End-to-End ASR
von: Jiang, Jintao, et al.
Veröffentlicht: (2024) -
Chain-of-Thought Reasoning in Streaming Full-Duplex End-to-End Spoken Dialogue Systems
von: Arora, Siddhant, et al.
Veröffentlicht: (2025)