Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Jain, Yash, Chan, David, Dheram, Pranav, Khare, Aparna, Shonibare, Olabanji, Ravichandran, Venkatesh, Ghosh, Shalini |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
por: Wang, Jinhan, et al.
Publicado: (2024)
por: Wang, Jinhan, et al.
Publicado: (2024)
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information
por: Taguchi, Chihiro, et al.
Publicado: (2024)
por: Taguchi, Chihiro, et al.
Publicado: (2024)
Towards Efficient Resume Understanding: A Multi-Granularity Multi-Modal Pre-Training Approach
por: Jiang, Feihu, et al.
Publicado: (2024)
por: Jiang, Feihu, et al.
Publicado: (2024)
Automatic Speech Recognition for Sanskrit with Transfer Learning
por: Sadhukhan, Bidit, et al.
Publicado: (2025)
por: Sadhukhan, Bidit, et al.
Publicado: (2025)
Automatic Speech Recognition for Greek Medical Dictation
por: Georgilas, Vardis, et al.
Publicado: (2025)
por: Georgilas, Vardis, et al.
Publicado: (2025)
Automatic Speech Recognition for Documenting Endangered Languages: Case Study of Ikema Miyakoan
por: Taguchi, Chihiro, et al.
Publicado: (2026)
por: Taguchi, Chihiro, et al.
Publicado: (2026)
Augmenting Automatic Speech Recognition Models with Disfluency Detection
por: Amann, Robin, et al.
Publicado: (2024)
por: Amann, Robin, et al.
Publicado: (2024)
Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio
por: He, Xinlu, et al.
Publicado: (2025)
por: He, Xinlu, et al.
Publicado: (2025)
Dual-Modality Multi-Stage Adversarial Safety Training: Robustifying Multimodal Web Agents Against Cross-Modal Attacks
por: Liu, Haoyu, et al.
Publicado: (2026)
por: Liu, Haoyu, et al.
Publicado: (2026)
InfiMed-Foundation: Pioneering Advanced Multimodal Medical Models with Compute-Efficient Pre-Training and Multi-Stage Fine-Tuning
por: Zhu, Guanghao, et al.
Publicado: (2025)
por: Zhu, Guanghao, et al.
Publicado: (2025)
Multi-Stage Verification-Centric Framework for Mitigating Hallucination in Multi-Modal RAG
por: Chen, Baiyu, et al.
Publicado: (2025)
por: Chen, Baiyu, et al.
Publicado: (2025)
Integrating Pre-Trained Speech and Language Models for End-to-End Speech Recognition
por: Hono, Yukiya, et al.
Publicado: (2023)
por: Hono, Yukiya, et al.
Publicado: (2023)
Multi-Modal Retrieval For Large Language Model Based Speech Recognition
por: Kolehmainen, Jari, et al.
Publicado: (2024)
por: Kolehmainen, Jari, et al.
Publicado: (2024)
Fairness of Automatic Speech Recognition: Looking Through a Philosophical Lens
por: Choi, Anna Seo Gyeong, et al.
Publicado: (2025)
por: Choi, Anna Seo Gyeong, et al.
Publicado: (2025)
Error-preserving Automatic Speech Recognition of Young English Learners' Language
por: Michot, Janick, et al.
Publicado: (2024)
por: Michot, Janick, et al.
Publicado: (2024)
Gujarati-English Code-Switching Speech Recognition using ensemble prediction of spoken language
por: Sharma, Yash, et al.
Publicado: (2024)
por: Sharma, Yash, et al.
Publicado: (2024)
LipGER: Visually-Conditioned Generative Error Correction for Robust Automatic Speech Recognition
por: Ghosh, Sreyan, et al.
Publicado: (2024)
por: Ghosh, Sreyan, et al.
Publicado: (2024)
Pre-Training Curriculum for Multi-Token Prediction in Language Models
por: Aynetdinov, Ansar, et al.
Publicado: (2025)
por: Aynetdinov, Ansar, et al.
Publicado: (2025)
Speech Retrieval-Augmented Generation without Automatic Speech Recognition
por: Min, Do June, et al.
Publicado: (2024)
por: Min, Do June, et al.
Publicado: (2024)
Handling Numeric Expressions in Automatic Speech Recognition
por: Huber, Christian, et al.
Publicado: (2024)
por: Huber, Christian, et al.
Publicado: (2024)
Doing More with Less: Data Augmentation for Sudanese Dialect Automatic Speech Recognition
por: Mansour, Ayman
Publicado: (2026)
por: Mansour, Ayman
Publicado: (2026)
A New Benchmark for Evaluating Automatic Speech Recognition in the Arabic Call Domain
por: Obaidah, Qusai Abo, et al.
Publicado: (2024)
por: Obaidah, Qusai Abo, et al.
Publicado: (2024)
MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference
por: Park, Jeonghyun, et al.
Publicado: (2025)
por: Park, Jeonghyun, et al.
Publicado: (2025)
Towards End-to-End Training of Automatic Speech Recognition for Nigerian Pidgin
por: Rufai, Amina Mardiyyah, et al.
Publicado: (2020)
por: Rufai, Amina Mardiyyah, et al.
Publicado: (2020)
SENS-ASR: Semantic Embedding injection in Neural-transducer for Streaming Automatic Speech Recognition
por: Dkhissi, Youness, et al.
Publicado: (2026)
por: Dkhissi, Youness, et al.
Publicado: (2026)
Benchmarking Rotary Position Embeddings for Automatic Speech Recognition
por: Zhang, Shucong, et al.
Publicado: (2025)
por: Zhang, Shucong, et al.
Publicado: (2025)
Embedding-to-Prefix: Parameter-Efficient Personalization for Pre-Trained Large Language Models
por: Huber, Bernd, et al.
Publicado: (2025)
por: Huber, Bernd, et al.
Publicado: (2025)
LoASR-Bench: Evaluating Large Speech Language Models on Low-Resource Automatic Speech Recognition Across Language Families
por: Chen, Jianan, et al.
Publicado: (2026)
por: Chen, Jianan, et al.
Publicado: (2026)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
por: Chan, David M., et al.
Publicado: (2024)
por: Chan, David M., et al.
Publicado: (2024)
Local Prompt Optimization
por: Jain, Yash, et al.
Publicado: (2025)
por: Jain, Yash, et al.
Publicado: (2025)
Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
por: Yang, Chao-Han Huck, et al.
Publicado: (2023)
por: Yang, Chao-Han Huck, et al.
Publicado: (2023)
Multi-Modal Automatic Prosody Annotation with Contrastive Pretraining of SSWP
por: Zhong, Jinzuomu, et al.
Publicado: (2023)
por: Zhong, Jinzuomu, et al.
Publicado: (2023)
Detecting and Mitigating Bias in LLMs through Knowledge Graph-Augmented Training
por: Kumar, Rajeev, et al.
Publicado: (2025)
por: Kumar, Rajeev, et al.
Publicado: (2025)
Semantically Corrected Amharic Automatic Speech Recognition
por: Adnew, Samuael, et al.
Publicado: (2024)
por: Adnew, Samuael, et al.
Publicado: (2024)
A Graph-based Approach for Multi-Modal Question Answering from Flowcharts in Telecom Documents
por: Soman, Sumit, et al.
Publicado: (2025)
por: Soman, Sumit, et al.
Publicado: (2025)
Rethinking Reflection in Pre-Training
por: AI, Essential, et al.
Publicado: (2025)
por: AI, Essential, et al.
Publicado: (2025)
Persona-Based Synthetic Data Generation Using Multi-Stage Conditioning with Large Language Models for Emotion Recognition
por: Inoshita, Keito, et al.
Publicado: (2025)
por: Inoshita, Keito, et al.
Publicado: (2025)
Learning to Trust the Crowd: A Multi-Model Consensus Reasoning Engine for Large Language Models
por: Kallem, Pranav
Publicado: (2026)
por: Kallem, Pranav
Publicado: (2026)
Improved Contextual Recognition In Automatic Speech Recognition Systems By Semantic Lattice Rescoring
por: Sudarshan, Ankitha, et al.
Publicado: (2023)
por: Sudarshan, Ankitha, et al.
Publicado: (2023)
MiMIC: Multi-Modal Indian Earnings Calls Dataset to Predict Stock Prices
por: Ghosh, Sohom, et al.
Publicado: (2025)
por: Ghosh, Sohom, et al.
Publicado: (2025)
Ejemplares similares
-
Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion
por: Wang, Jinhan, et al.
Publicado: (2024) -
Killkan: The Automatic Speech Recognition Dataset for Kichwa with Morphosyntactic Information
por: Taguchi, Chihiro, et al.
Publicado: (2024) -
Towards Efficient Resume Understanding: A Multi-Granularity Multi-Modal Pre-Training Approach
por: Jiang, Feihu, et al.
Publicado: (2024) -
Automatic Speech Recognition for Sanskrit with Transfer Learning
por: Sadhukhan, Bidit, et al.
Publicado: (2025) -
Automatic Speech Recognition for Greek Medical Dictation
por: Georgilas, Vardis, et al.
Publicado: (2025)