End-to-end Automatic Speech Recognition and Speech Translation: Integration of Speech Foundational Models and LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Luu, Nam, Bojar, Ondřej |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
von: Polák, Peter, et al.
Veröffentlicht: (2023)
von: Polák, Peter, et al.
Veröffentlicht: (2023)
Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation
von: Javorský, Dávid, et al.
Veröffentlicht: (2022)
von: Javorský, Dávid, et al.
Veröffentlicht: (2022)
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
von: Stankov, Vladislav, et al.
Veröffentlicht: (2025)
von: Stankov, Vladislav, et al.
Veröffentlicht: (2025)
Quality and Quantity of Machine Translation References for Automatic Metrics
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024)
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
von: Polák, Peter, et al.
Veröffentlicht: (2025)
von: Polák, Peter, et al.
Veröffentlicht: (2025)
How "Real" is Your Real-Time Simultaneous Speech-to-Text Translation System?
von: Papi, Sara, et al.
Veröffentlicht: (2024)
von: Papi, Sara, et al.
Veröffentlicht: (2024)
Evaluating the IWSLT2023 Speech Translation Tasks: Human Annotations, Automatic Metrics, and Segmentation
von: Sperber, Matthias, et al.
Veröffentlicht: (2024)
von: Sperber, Matthias, et al.
Veröffentlicht: (2024)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
von: Shakeel, Muhammad, et al.
Veröffentlicht: (2024)
ViSpeechFormer: A Phonemic Approach for Vietnamese Automatic Speech Recognition
von: Nguyen, Khoa Anh, et al.
Veröffentlicht: (2026)
von: Nguyen, Khoa Anh, et al.
Veröffentlicht: (2026)
End-to-end Speech Recognition with similar length speech and text
von: Fan, Peng, et al.
Veröffentlicht: (2025)
von: Fan, Peng, et al.
Veröffentlicht: (2025)
Prompting LLMs: Length Control for Isometric Machine Translation
von: Javorský, Dávid, et al.
Veröffentlicht: (2025)
von: Javorský, Dávid, et al.
Veröffentlicht: (2025)
Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition
von: Rong, Yiming, et al.
Veröffentlicht: (2025)
von: Rong, Yiming, et al.
Veröffentlicht: (2025)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
von: Hono, Yukiya, et al.
Veröffentlicht: (2023)
von: Hono, Yukiya, et al.
Veröffentlicht: (2023)
Syllabic-Structure Decoder for Automatic Speech Recognition in Vietnamese
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2026)
von: Nguyen, Nghia Hieu, et al.
Veröffentlicht: (2026)
WildSpeech-Bench: Benchmarking End-to-End SpeechLLMs in the Wild
von: Zhang, Linhao, et al.
Veröffentlicht: (2025)
von: Zhang, Linhao, et al.
Veröffentlicht: (2025)
Data Augmentation for End-to-end Code-switching Speech Recognition
von: Du, Chenpeng, et al.
Veröffentlicht: (2020)
von: Du, Chenpeng, et al.
Veröffentlicht: (2020)
Decoder-only Architecture for Streaming End-to-end Speech Recognition
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2024)
Continual Learning for Monolingual End-to-End Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2021)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2021)
Zero-resource Speech Translation and Recognition with LLMs
von: Mundnich, Karel, et al.
Veröffentlicht: (2024)
von: Mundnich, Karel, et al.
Veröffentlicht: (2024)
LLaST: Improved End-to-end Speech Translation System Leveraged by Large Language Models
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
Speech Translation with Speech Foundation Models and Large Language Models: What is There and What is Missing?
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
Understanding the role of FFNs in driving multilingual behaviour in LLMs
von: Bhattacharya, Sunit, et al.
Veröffentlicht: (2024)
von: Bhattacharya, Sunit, et al.
Veröffentlicht: (2024)
Finetuning LLMs for EvaCun 2025 token prediction shared task
von: Jon, Josef, et al.
Veröffentlicht: (2025)
von: Jon, Josef, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition for the Ika Language
von: Nzenwata, Uchenna, et al.
Veröffentlicht: (2024)
von: Nzenwata, Uchenna, et al.
Veröffentlicht: (2024)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
von: Peng, Yifan, et al.
Veröffentlicht: (2024)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
HITSZ's End-To-End Speech Translation Systems Combining Sequence-to-Sequence Auto Speech Recognition Model and Indic Large Language Model for IWSLT 2025 in Indic Track
von: Wei, Xuchen, et al.
Veröffentlicht: (2025)
von: Wei, Xuchen, et al.
Veröffentlicht: (2025)
Word Level Timestamp Generation for Automatic Speech Recognition and Translation
von: Hu, Ke, et al.
Veröffentlicht: (2025)
von: Hu, Ke, et al.
Veröffentlicht: (2025)
PROST-LLM: Progressively Enhancing the Speech-to-Speech Translation Capability in LLMs
von: Xu, Jing, et al.
Veröffentlicht: (2026)
von: Xu, Jing, et al.
Veröffentlicht: (2026)
Automatic Speech Recognition for Hindi
von: Saha, Anish, et al.
Veröffentlicht: (2024)
von: Saha, Anish, et al.
Veröffentlicht: (2024)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
von: Rufai, Amina Mardiyyah, et al.
Veröffentlicht: (2020)
Representation Purification for End-to-End Speech Translation
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
von: Zhang, Chengwei, et al.
Veröffentlicht: (2024)
A Case Study on Filtering for End-to-End Speech Translation
von: Alam, Md Mahfuz Ibn, et al.
Veröffentlicht: (2024)
von: Alam, Md Mahfuz Ibn, et al.
Veröffentlicht: (2024)
Pushing the Limits of Zero-shot End-to-End Speech Translation
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Responsible Benchmarking of Fairness for Automatic Speech Recognition
von: Herron, Felix, et al.
Veröffentlicht: (2026)
von: Herron, Felix, et al.
Veröffentlicht: (2026)
Vietnamese Automatic Speech Recognition: A Revisit
von: Vu, Thi, et al.
Veröffentlicht: (2026)
von: Vu, Thi, et al.
Veröffentlicht: (2026)
Hearing to Translate: The Effectiveness of Speech Modality Integration into LLMs
von: Papi, Sara, et al.
Veröffentlicht: (2025)
von: Papi, Sara, et al.
Veröffentlicht: (2025)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
von: He, Xinlu, et al.
Veröffentlicht: (2025)
von: He, Xinlu, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Long-Form End-to-End Speech Translation via Latent Alignment Segmentation
von: Polák, Peter, et al.
Veröffentlicht: (2023) -
Continuous Rating as Reliable Human Evaluation of Simultaneous Speech Translation
von: Javorský, Dávid, et al.
Veröffentlicht: (2022) -
ParCzech4Speech: A New Speech Corpus Derived from Czech Parliamentary Data
von: Stankov, Vladislav, et al.
Veröffentlicht: (2025) -
Quality and Quantity of Machine Translation References for Automatic Metrics
von: Zouhar, Vilém, et al.
Veröffentlicht: (2024) -
Better Late Than Never: Meta-Evaluation of Latency Metrics for Simultaneous Speech-to-Text Translation
von: Polák, Peter, et al.
Veröffentlicht: (2025)