Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
Fuente:
arXiv
Guardado en:
| Autores principales: | Fernandez-Lopez, Adriana, Liu, Shiwei, Yin, Lu, Petridis, Stavros, Pantic, Maja |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dynamic Data Pruning for Automatic Speech Recognition
por: Xiao, Qiao, et al.
Publicado: (2024)
por: Xiao, Qiao, et al.
Publicado: (2024)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
por: Cappellazzo, Umberto, et al.
Publicado: (2026)
por: Cappellazzo, Umberto, et al.
Publicado: (2026)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
por: Anand, et al.
Publicado: (2025)
por: Anand, et al.
Publicado: (2025)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
por: Chen, Honglie, et al.
Publicado: (2024)
por: Chen, Honglie, et al.
Publicado: (2024)
Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
por: Piñeiro-Martín, Andrés, et al.
Publicado: (2024)
por: Piñeiro-Martín, Andrés, et al.
Publicado: (2024)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
por: Kim, Minsu, et al.
Publicado: (2025)
por: Kim, Minsu, et al.
Publicado: (2025)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
por: Zhao, Qiuming, et al.
Publicado: (2025)
por: Zhao, Qiuming, et al.
Publicado: (2025)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
por: Ugan, Enes Yavuz, et al.
Publicado: (2025)
Sequential Editing for Lifelong Training of Speech Recognition Models
por: Kulshreshtha, Devang, et al.
Publicado: (2024)
por: Kulshreshtha, Devang, et al.
Publicado: (2024)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
por: Dao, Alan, et al.
Publicado: (2025)
por: Dao, Alan, et al.
Publicado: (2025)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
por: Hsu, Ming-Hao, et al.
Publicado: (2024)
por: Hsu, Ming-Hao, et al.
Publicado: (2024)
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
por: Wang, Yiming, et al.
Publicado: (2023)
por: Wang, Yiming, et al.
Publicado: (2023)
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages
por: Ayrapetyan, Alexan, et al.
Publicado: (2025)
por: Ayrapetyan, Alexan, et al.
Publicado: (2025)
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
por: Shen, Peng, et al.
Publicado: (2025)
por: Shen, Peng, et al.
Publicado: (2025)
Automatic Speech Recognition for African Low-Resource Languages: Challenges and Future Directions
por: Imam, Sukairaj Hafiz, et al.
Publicado: (2025)
por: Imam, Sukairaj Hafiz, et al.
Publicado: (2025)
Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2022)
por: Eeckt, Steven Vander, et al.
Publicado: (2022)
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
por: Wang, Tianduo, et al.
Publicado: (2025)
por: Wang, Tianduo, et al.
Publicado: (2025)
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
por: Yang, Zhengdong, et al.
Publicado: (2025)
por: Yang, Zhengdong, et al.
Publicado: (2025)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
por: Wang, Haoyu, et al.
Publicado: (2022)
por: Wang, Haoyu, et al.
Publicado: (2022)
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
por: Naouara, Hedi, et al.
Publicado: (2025)
por: Naouara, Hedi, et al.
Publicado: (2025)
Automatic Speech Recognition for Hindi
por: Saha, Anish, et al.
Publicado: (2024)
por: Saha, Anish, et al.
Publicado: (2024)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
por: Tsunoo, Emiru, et al.
Publicado: (2025)
por: Tsunoo, Emiru, et al.
Publicado: (2025)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
por: Li, Chin-Jou, et al.
Publicado: (2025)
por: Li, Chin-Jou, et al.
Publicado: (2025)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
por: Casanova, Edresson, et al.
Publicado: (2024)
por: Casanova, Edresson, et al.
Publicado: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
por: Tsiamas, Ioannis, et al.
Publicado: (2024)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
por: Xu, Tianyi, et al.
Publicado: (2025)
por: Xu, Tianyi, et al.
Publicado: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2024)
por: Shivakumar, Prashanth Gurunath, et al.
Publicado: (2024)
Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages
por: Xiao, Yang, et al.
Publicado: (2026)
por: Xiao, Yang, et al.
Publicado: (2026)
EmoTale: An Enacted Speech-emotion Dataset in Danish
por: Hjuler, Maja J., et al.
Publicado: (2025)
por: Hjuler, Maja J., et al.
Publicado: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
por: Futami, Hayato, et al.
Publicado: (2025)
por: Futami, Hayato, et al.
Publicado: (2025)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
por: Hu, Jiliang, et al.
Publicado: (2025)
por: Hu, Jiliang, et al.
Publicado: (2025)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
por: Wei, Kun, et al.
Publicado: (2023)
por: Wei, Kun, et al.
Publicado: (2023)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
por: Vesterbacka, Leonora, et al.
Publicado: (2025)
por: Vesterbacka, Leonora, et al.
Publicado: (2025)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
por: Lee, Jeehyun, et al.
Publicado: (2024)
por: Lee, Jeehyun, et al.
Publicado: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
por: Wang, Yujin, et al.
Publicado: (2022)
por: Wang, Yujin, et al.
Publicado: (2022)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
por: Cornell, Samuele, et al.
Publicado: (2024)
por: Cornell, Samuele, et al.
Publicado: (2024)
From Statistical Methods to Pre-Trained Models; A Survey on Automatic Speech Recognition for Resource Scarce Urdu Language
por: Sharif, Muhammad, et al.
Publicado: (2024)
por: Sharif, Muhammad, et al.
Publicado: (2024)
Ejemplares similares
-
Dynamic Data Pruning for Automatic Speech Recognition
por: Xiao, Qiao, et al.
Publicado: (2024) -
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
por: Cappellazzo, Umberto, et al.
Publicado: (2026) -
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
por: Anand, et al.
Publicado: (2025) -
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
por: Cappellazzo, Umberto, et al.
Publicado: (2025) -
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
por: Chen, Honglie, et al.
Publicado: (2024)