Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fernandez-Lopez, Adriana, Liu, Shiwei, Yin, Lu, Petridis, Stavros, Pantic, Maja |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Data Pruning for Automatic Speech Recognition
von: Xiao, Qiao, et al.
Veröffentlicht: (2024)
von: Xiao, Qiao, et al.
Veröffentlicht: (2024)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
von: Anand, et al.
Veröffentlicht: (2025)
von: Anand, et al.
Veröffentlicht: (2025)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
von: Chen, Honglie, et al.
Veröffentlicht: (2024)
von: Chen, Honglie, et al.
Veröffentlicht: (2024)
Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
von: Piñeiro-Martín, Andrés, et al.
Veröffentlicht: (2024)
von: Piñeiro-Martín, Andrés, et al.
Veröffentlicht: (2024)
MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025)
Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
von: Kim, Minsu, et al.
Veröffentlicht: (2025)
Low-Rank and Sparse Model Merging for Multi-Lingual Speech Recognition and Translation
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025)
von: Zhao, Qiuming, et al.
Veröffentlicht: (2025)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2024)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
von: Ugan, Enes Yavuz, et al.
Veröffentlicht: (2025)
Sequential Editing for Lifelong Training of Speech Recognition Models
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
von: Kulshreshtha, Devang, et al.
Veröffentlicht: (2024)
Speechless: Speech Instruction Training Without Speech for Low Resource Languages
von: Dao, Alan, et al.
Veröffentlicht: (2025)
von: Dao, Alan, et al.
Veröffentlicht: (2025)
SMILE: Speech Meta In-Context Learning for Low-Resource Language Automatic Speech Recognition
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
von: Hsu, Ming-Hao, et al.
Veröffentlicht: (2024)
ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
von: Wang, Yiming, et al.
Veröffentlicht: (2023)
von: Wang, Yiming, et al.
Veröffentlicht: (2023)
Methods to Increase the Amount of Data for Speech Recognition for Low Resource Languages
von: Ayrapetyan, Alexan, et al.
Veröffentlicht: (2025)
von: Ayrapetyan, Alexan, et al.
Veröffentlicht: (2025)
Retrieval-Augmented Speech Recognition Approach for Domain Challenges
von: Shen, Peng, et al.
Veröffentlicht: (2025)
von: Shen, Peng, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition for African Low-Resource Languages: Challenges and Future Directions
von: Imam, Sukairaj Hafiz, et al.
Veröffentlicht: (2025)
von: Imam, Sukairaj Hafiz, et al.
Veröffentlicht: (2025)
Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)
von: Eeckt, Steven Vander, et al.
Veröffentlicht: (2022)
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition
von: Wang, Tianduo, et al.
Veröffentlicht: (2025)
von: Wang, Tianduo, et al.
Veröffentlicht: (2025)
Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
von: Yang, Zhengdong, et al.
Veröffentlicht: (2025)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
LinTO Audio and Textual Datasets to Train and Evaluate Automatic Speech Recognition in Tunisian Arabic Dialect
von: Naouara, Hedi, et al.
Veröffentlicht: (2025)
von: Naouara, Hedi, et al.
Veröffentlicht: (2025)
Automatic Speech Recognition for Hindi
von: Saha, Anish, et al.
Veröffentlicht: (2024)
von: Saha, Anish, et al.
Veröffentlicht: (2024)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025)
von: Tsunoo, Emiru, et al.
Veröffentlicht: (2025)
Towards Inclusive ASR: Investigating Voice Conversion for Dysarthric Speech Recognition in Low-Resource Languages
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
von: Li, Chin-Jou, et al.
Veröffentlicht: (2025)
Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
von: Casanova, Edresson, et al.
Veröffentlicht: (2024)
Speech is More Than Words: Do Speech-to-Text Translation Systems Leverage Prosody?
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
von: Tsiamas, Ioannis, et al.
Veröffentlicht: (2024)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
Speech Recognition Rescoring with Large Speech-Text Foundation Models
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
von: Shivakumar, Prashanth Gurunath, et al.
Veröffentlicht: (2024)
Adapting Where It Matters: Depth-Aware Adaptation for Efficient Multilingual Speech Recognition in Low-Resource Languages
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
von: Xiao, Yang, et al.
Veröffentlicht: (2026)
EmoTale: An Enacted Speech-emotion Dataset in Danish
von: Hjuler, Maja J., et al.
Veröffentlicht: (2025)
von: Hjuler, Maja J., et al.
Veröffentlicht: (2025)
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
von: Futami, Hayato, et al.
Veröffentlicht: (2025)
Joint Automatic Speech Recognition And Structure Learning For Better Speech Understanding
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
von: Hu, Jiliang, et al.
Veröffentlicht: (2025)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
von: Wei, Kun, et al.
Veröffentlicht: (2023)
von: Wei, Kun, et al.
Veröffentlicht: (2023)
Swedish Whispers; Leveraging a Massive Speech Corpus for Swedish Speech Recognition
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
von: Vesterbacka, Leonora, et al.
Veröffentlicht: (2025)
Inappropriate Pause Detection In Dysarthric Speech Using Large-Scale Speech Recognition
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
von: Lee, Jeehyun, et al.
Veröffentlicht: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
von: Cornell, Samuele, et al.
Veröffentlicht: (2024)
From Statistical Methods to Pre-Trained Models; A Survey on Automatic Speech Recognition for Resource Scarce Urdu Language
von: Sharif, Muhammad, et al.
Veröffentlicht: (2024)
von: Sharif, Muhammad, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dynamic Data Pruning for Automatic Speech Recognition
von: Xiao, Qiao, et al.
Veröffentlicht: (2024) -
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2026) -
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
von: Anand, et al.
Veröffentlicht: (2025) -
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
von: Cappellazzo, Umberto, et al.
Veröffentlicht: (2025) -
RT-LA-VocE: Real-Time Low-SNR Audio-Visual Speech Enhancement
von: Chen, Honglie, et al.
Veröffentlicht: (2024)