Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
Fuente:
arXiv
Guardado en:
| Autores principales: | Wright, George August, Cappellazzo, Umberto, Zaiem, Salah, Raj, Desh, Yang, Lucas Ondel, Falavigna, Daniele, Ali, Mohamed Nabih, Brutti, Alessio |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
por: Lasbordes, Maxence, et al.
Publicado: (2025)
por: Lasbordes, Maxence, et al.
Publicado: (2025)
Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers
por: Cappellazzo, Umberto, et al.
Publicado: (2023)
por: Cappellazzo, Umberto, et al.
Publicado: (2023)
Continual Contrastive Spoken Language Understanding
por: Cappellazzo, Umberto, et al.
Publicado: (2023)
por: Cappellazzo, Umberto, et al.
Publicado: (2023)
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
Input Conditioned Layer Dropping in Speech Foundation Models
por: Hannan, Abdul, et al.
Publicado: (2025)
por: Hannan, Abdul, et al.
Publicado: (2025)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
por: Cappellazzo, Umberto, et al.
Publicado: (2024)
MLMA: Towards Multilingual ASR With Mamba-based Architectures
por: Ali, Mohamed Nabih, et al.
Publicado: (2025)
por: Ali, Mohamed Nabih, et al.
Publicado: (2025)
Federating Dynamic Models using Early-Exit Architectures for Automatic Speech Recognition on Heterogeneous Clients
por: Ali, Mohamed Nabih, et al.
Publicado: (2024)
por: Ali, Mohamed Nabih, et al.
Publicado: (2024)
Listening to Multi-talker Conversations: Modular and End-to-end Perspectives
por: Raj, Desh
Publicado: (2024)
por: Raj, Desh
Publicado: (2024)
Prominence-aware automatic speech recognition for conversational speech
por: Linke, Julian, et al.
Publicado: (2025)
por: Linke, Julian, et al.
Publicado: (2025)
Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations
por: Zaiem, Salah, et al.
Publicado: (2024)
por: Zaiem, Salah, et al.
Publicado: (2024)
Zipformer: A faster and better encoder for automatic speech recognition
por: Yao, Zengwei, et al.
Publicado: (2023)
por: Yao, Zengwei, et al.
Publicado: (2023)
Robustifying automatic speech recognition by extracting slowly varying features
por: Pizarro, Matías, et al.
Publicado: (2021)
por: Pizarro, Matías, et al.
Publicado: (2021)
AS-70: A Mandarin stuttered speech dataset for automatic speech recognition and stuttering event detection
por: Gong, Rong, et al.
Publicado: (2024)
por: Gong, Rong, et al.
Publicado: (2024)
Evaluating and Improving Continual Learning in Spoken Language Understanding
por: Yang, Muqiao, et al.
Publicado: (2024)
por: Yang, Muqiao, et al.
Publicado: (2024)
Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition
por: Cappellazzo, Umberto, et al.
Publicado: (2026)
por: Cappellazzo, Umberto, et al.
Publicado: (2026)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
por: Zhang, Yiru, et al.
Publicado: (2025)
por: Zhang, Yiru, et al.
Publicado: (2025)
Adaptive Audio-Visual Speech Recognition via Matryoshka-Based Multimodal LLMs
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
The evaluation of a code-switched Sepedi-English automatic speech recognition system
por: Phaladi, Amanda, et al.
Publicado: (2024)
por: Phaladi, Amanda, et al.
Publicado: (2024)
SeMaScore : a new evaluation metric for automatic speech recognition tasks
por: Sasindran, Zitha, et al.
Publicado: (2024)
por: Sasindran, Zitha, et al.
Publicado: (2024)
Automated speech audiometry: Can it work using open-source pre-trained Kaldi-NL automatic speech recognition?
por: Araiza-Illan, Gloria, et al.
Publicado: (2023)
por: Araiza-Illan, Gloria, et al.
Publicado: (2023)
Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs
por: Anand, et al.
Publicado: (2025)
por: Anand, et al.
Publicado: (2025)
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
por: Gaido, Marco, et al.
Publicado: (2024)
por: Gaido, Marco, et al.
Publicado: (2024)
Speech Self-Supervised Representations Benchmarking: a Case for Larger Probing Heads
por: Zaiem, Salah, et al.
Publicado: (2023)
por: Zaiem, Salah, et al.
Publicado: (2023)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
por: San, Nay, et al.
Publicado: (2024)
por: San, Nay, et al.
Publicado: (2024)
Lightweight End-to-end Text-to-speech Synthesis for low resource on-device applications
por: Vecino, Biel Tura, et al.
Publicado: (2025)
por: Vecino, Biel Tura, et al.
Publicado: (2025)
CognoSpeak: an automatic, remote assessment of early cognitive decline in real-world conversational speech
por: Pahar, Madhurananda, et al.
Publicado: (2025)
por: Pahar, Madhurananda, et al.
Publicado: (2025)
Faster Speech-LLaMA Inference with Multi-token Prediction
por: Raj, Desh, et al.
Publicado: (2024)
por: Raj, Desh, et al.
Publicado: (2024)
Introduction to speech recognition
por: Dauphin, Gabriel
Publicado: (2024)
por: Dauphin, Gabriel
Publicado: (2024)
Advancing automatic speech recognition using feature fusion with self-supervised learning features: A case study on Fearless Steps Apollo corpus
por: Chen, Szu-Jui, et al.
Publicado: (2026)
por: Chen, Szu-Jui, et al.
Publicado: (2026)
Improving child speech recognition with augmented child-like speech
por: Zhang, Yuanyuan, et al.
Publicado: (2024)
por: Zhang, Yuanyuan, et al.
Publicado: (2024)
Low-resource speech recognition and dialect identification of Irish in a multi-task framework
por: Lonergan, Liam, et al.
Publicado: (2024)
por: Lonergan, Liam, et al.
Publicado: (2024)
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
por: Fong, Seraphina, et al.
Publicado: (2025)
por: Fong, Seraphina, et al.
Publicado: (2025)
Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
por: Cappellazzo, Umberto, et al.
Publicado: (2025)
Robust fine-tuning of speech recognition models via model merging: application to disordered speech
por: Ducorroy, Alexandre, et al.
Publicado: (2025)
por: Ducorroy, Alexandre, et al.
Publicado: (2025)
Perceptual implications of automatic anonymization in pathological speech
por: Arasteh, Soroosh Tayebi, et al.
Publicado: (2025)
por: Arasteh, Soroosh Tayebi, et al.
Publicado: (2025)
An automatic mixing speech enhancement system for multi-track audio
por: Liu, Xiaojing, et al.
Publicado: (2024)
por: Liu, Xiaojing, et al.
Publicado: (2024)
Teaching the Teachers: Boosting unsupervised domain adaptation in speech recognition by ensemble update
por: Ahmad, Rehan, et al.
Publicado: (2026)
por: Ahmad, Rehan, et al.
Publicado: (2026)
Joint decoding method for controllable contextual speech recognition based on Speech LLM
por: Fang, Yangui, et al.
Publicado: (2025)
por: Fang, Yangui, et al.
Publicado: (2025)
Ejemplares similares
-
Splitformer: An improved early-exit architecture for automatic speech recognition on edge devices
por: Lasbordes, Maxence, et al.
Publicado: (2025) -
Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters
por: Cappellazzo, Umberto, et al.
Publicado: (2024) -
Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers
por: Cappellazzo, Umberto, et al.
Publicado: (2023) -
Continual Contrastive Spoken Language Understanding
por: Cappellazzo, Umberto, et al.
Publicado: (2023) -
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach
por: Cappellazzo, Umberto, et al.
Publicado: (2025)