Bridging the gap: A comparative exploration of Speech-LLM and end-to-end architecture for multilingual conversational ASR
Fuente:
arXiv
Guardado en:
| Autores principales: | Mei, Yuxiang, Xu, Dongxing, Liang, Jiaen, Long, Yanhua |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
por: Zheng, Yuang, et al.
Publicado: (2026)
por: Zheng, Yuang, et al.
Publicado: (2026)
Improving endpoint detection in end-to-end streaming ASR for conversational speech
por: C, Anandh, et al.
Publicado: (2025)
por: C, Anandh, et al.
Publicado: (2025)
SHNU Multilingual Conversational Speech Recognition System for INTERSPEECH 2025 MLC-SLM Challenge
por: Mei, Yuxiang, et al.
Publicado: (2025)
por: Mei, Yuxiang, et al.
Publicado: (2025)
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
por: Cheng, Shanbo, et al.
Publicado: (2024)
por: Cheng, Shanbo, et al.
Publicado: (2024)
Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
por: Chen, Shuangyuan, et al.
Publicado: (2025)
por: Chen, Shuangyuan, et al.
Publicado: (2025)
End-to-end Joint Punctuated and Normalized ASR with a Limited Amount of Punctuated Training Data
por: Cui, Can, et al.
Publicado: (2023)
por: Cui, Can, et al.
Publicado: (2023)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
por: Cheng, Shanbo, et al.
Publicado: (2025)
por: Cheng, Shanbo, et al.
Publicado: (2025)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
por: Žavoronkov, Aleksei, et al.
Publicado: (2025)
por: Žavoronkov, Aleksei, et al.
Publicado: (2025)
Text-only domain adaptation for end-to-end ASR using integrated text-to-mel-spectrogram generator
por: Bataev, Vladimir, et al.
Publicado: (2023)
por: Bataev, Vladimir, et al.
Publicado: (2023)
Bridging Speech and Text: Enhancing ASR with Pinyin-to-Character Pre-training in LLMs
por: Yuhang, Yang, et al.
Publicado: (2024)
por: Yuhang, Yang, et al.
Publicado: (2024)
A two-stage transliteration approach to improve performance of a multilingual ASR
por: Kumar, Rohit
Publicado: (2024)
por: Kumar, Rohit
Publicado: (2024)
Keep Decoding Parallel with Effective Knowledge Distillation from Language Models to End-to-end Speech Recognisers
por: Hentschel, Michael, et al.
Publicado: (2024)
por: Hentschel, Michael, et al.
Publicado: (2024)
ICSD: An Open-source Dataset for Infant Cry and Snoring Detection
por: Liu, Qingyu, et al.
Publicado: (2024)
por: Liu, Qingyu, et al.
Publicado: (2024)
Rethinking Entropy Allocation in LLM-based ASR: Understanding the Dynamics between Speech Encoders and LLMs
por: Xie, Yuan, et al.
Publicado: (2026)
por: Xie, Yuan, et al.
Publicado: (2026)
Configurable Multilingual ASR with Speech Summary Representations
por: Zhu, Harrison, et al.
Publicado: (2024)
por: Zhu, Harrison, et al.
Publicado: (2024)
ContextASR-Bench: A Massive Contextual Speech Recognition Benchmark
por: Wang, He, et al.
Publicado: (2025)
por: Wang, He, et al.
Publicado: (2025)
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
por: Lin, Yi-Cheng, et al.
Publicado: (2026)
por: Lin, Yi-Cheng, et al.
Publicado: (2026)
Benchmarking Japanese Speech Recognition on ASR-LLM Setups with Multi-Pass Augmented Generative Error Correction
por: Ko, Yuka, et al.
Publicado: (2024)
por: Ko, Yuka, et al.
Publicado: (2024)
Performant ASR Models for Medical Entities in Accented Speech
por: Afonja, Tejumade, et al.
Publicado: (2024)
por: Afonja, Tejumade, et al.
Publicado: (2024)
Crossmodal ASR Error Correction with Discrete Speech Units
por: Li, Yuanchao, et al.
Publicado: (2024)
por: Li, Yuanchao, et al.
Publicado: (2024)
PromptASR for contextualized ASR with controllable style
por: Yang, Xiaoyu, et al.
Publicado: (2023)
por: Yang, Xiaoyu, et al.
Publicado: (2023)
Advancing LLM-based phoneme-to-grapheme for multilingual speech recognition
por: Dong, Lukuang, et al.
Publicado: (2026)
por: Dong, Lukuang, et al.
Publicado: (2026)
NIM4-ASR: Towards Efficient, Robust, and Customizable Real-Time LLM-Based ASR
por: Xie, Yuan, et al.
Publicado: (2026)
por: Xie, Yuan, et al.
Publicado: (2026)
Exploring SSL Discrete Speech Features for Zipformer-based Contextual ASR
por: Cui, Mingyu, et al.
Publicado: (2024)
por: Cui, Mingyu, et al.
Publicado: (2024)
Interventional Speech Noise Injection for ASR Generalizable Spoken Language Understanding
por: Jung, Yeonjoon, et al.
Publicado: (2024)
por: Jung, Yeonjoon, et al.
Publicado: (2024)
ASR-FAIRBENCH: Measuring and Benchmarking Equity Across Speech Recognition Systems
por: Rai, Anand, et al.
Publicado: (2025)
por: Rai, Anand, et al.
Publicado: (2025)
AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR
por: Nguyen, Tuan, et al.
Publicado: (2025)
por: Nguyen, Tuan, et al.
Publicado: (2025)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
por: Fan, Ruchao, et al.
Publicado: (2024)
por: Fan, Ruchao, et al.
Publicado: (2024)
Unveiling the Potential of LLM-Based ASR on Chinese Open-Source Datasets
por: Geng, Xuelong, et al.
Publicado: (2024)
por: Geng, Xuelong, et al.
Publicado: (2024)
Performance Analysis of Speech Encoders for Low-Resource SLU and ASR in Tunisian Dialect
por: Mdhaffar, Salima, et al.
Publicado: (2024)
por: Mdhaffar, Salima, et al.
Publicado: (2024)
LibriheavyMix: A 20,000-Hour Dataset for Single-Channel Reverberant Multi-Talker Speech Separation, ASR and Speaker Diarization
por: Jin, Zengrui, et al.
Publicado: (2024)
por: Jin, Zengrui, et al.
Publicado: (2024)
UniEnc-CASSNAT: An Encoder-only Non-autoregressive ASR for Speech SSL Models
por: Fan, Ruchao, et al.
Publicado: (2024)
por: Fan, Ruchao, et al.
Publicado: (2024)
Advancing Hearing Assessment: An ASR-Based Frequency-Specific Speech Test for Diagnosing Presbycusis
por: Bleeck, Stefan
Publicado: (2025)
por: Bleeck, Stefan
Publicado: (2025)
ASR Under the Stethoscope: Evaluating Biases in Clinical Speech Recognition across Indian Languages
por: Kumar, Subham, et al.
Publicado: (2025)
por: Kumar, Subham, et al.
Publicado: (2025)
A Comprehensive Study on the Effectiveness of ASR Representations for Noise-Robust Speech Emotion Recognition
por: Shi, Xiaohan, et al.
Publicado: (2023)
por: Shi, Xiaohan, et al.
Publicado: (2023)
TokenVerse: Towards Unifying Speech and NLP Tasks via Transducer-based ASR
por: Kumar, Shashi, et al.
Publicado: (2024)
por: Kumar, Shashi, et al.
Publicado: (2024)
Echotune: A Modular Extractor Leveraging the Variable-Length Nature of Speech in ASR Tasks
por: Chen, Sizhou, et al.
Publicado: (2023)
por: Chen, Sizhou, et al.
Publicado: (2023)
Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades
por: Jung, Donghyuk, et al.
Publicado: (2026)
por: Jung, Donghyuk, et al.
Publicado: (2026)
Evolutionary Prompt Design for LLM-Based Post-ASR Error Correction
por: Sachdev, Rithik, et al.
Publicado: (2024)
por: Sachdev, Rithik, et al.
Publicado: (2024)
Ejemplares similares
-
A Language-Agnostic Hierarchical LoRA-MoE Architecture for CTC-based Multilingual ASR
por: Zheng, Yuang, et al.
Publicado: (2026) -
Improving endpoint detection in end-to-end streaming ASR for conversational speech
por: C, Anandh, et al.
Publicado: (2025) -
SHNU Multilingual Conversational Speech Recognition System for INTERSPEECH 2025 MLC-SLM Challenge
por: Mei, Yuxiang, et al.
Publicado: (2025) -
Towards Achieving Human Parity on End-to-end Simultaneous Speech Translation via LLM Agent
por: Cheng, Shanbo, et al.
Publicado: (2024) -
Noisy Disentanglement with Tri-stage Training for Noise-Robust Speech Recognition
por: Chen, Shuangyuan, et al.
Publicado: (2025)