An Effective Training Framework for Light-Weight Automatic Speech Recognition Models
Fuente:
arXiv
Saved in:
| Main Authors: | Hannan, Abdul, Brutti, Alessio, Nawaz, Shah, Noman, Mubashir |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Distillation-based Layer Dropping (DLD): Effective End-to-end Framework for Dynamic Speech Networks
by: Hannan, Abdul, et al.
Published: (2026)
by: Hannan, Abdul, et al.
Published: (2026)
Input Conditioned Layer Dropping in Speech Foundation Models
by: Hannan, Abdul, et al.
Published: (2025)
by: Hannan, Abdul, et al.
Published: (2025)
EoCD: Encoder only Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2026)
by: Noman, Mubashir, et al.
Published: (2026)
PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association
by: Hannan, Abdul, et al.
Published: (2025)
by: Hannan, Abdul, et al.
Published: (2025)
RFOP: Rethinking Fusion and Orthogonal Projection for Face-Voice Association
by: Hannan, Abdul, et al.
Published: (2025)
by: Hannan, Abdul, et al.
Published: (2025)
InceptionMamba: Efficient Multi-Stage Feature Enhancement with Selective State Space Model for Microscopic Medical Image Segmentation
by: Kareem, Daniya Najiha Abdul, et al.
Published: (2025)
by: Kareem, Daniya Najiha Abdul, et al.
Published: (2025)
SB-BEVFusion: Enhancing the Robustness against Sensor Malfunction and Corruptions
by: Essl, Markus, et al.
Published: (2026)
by: Essl, Markus, et al.
Published: (2026)
Face-Voice Association with Inductive Bias for Maximum Class Separation
by: Moscati, Marta, et al.
Published: (2026)
by: Moscati, Marta, et al.
Published: (2026)
ChangeBind: A Hybrid Change Encoder for Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
HyRet-Change: A hybrid retentive network for remote sensing change detection
by: Fiaz, Mustansar, et al.
Published: (2025)
by: Fiaz, Mustansar, et al.
Published: (2025)
FANet: Feature Amplification Network for Semantic Segmentation in Cluttered Background
by: Ali, Muhammad, et al.
Published: (2024)
by: Ali, Muhammad, et al.
Published: (2024)
COSNet: A Novel Semantic Segmentation Network using Enhanced Boundaries in Cluttered Scenes
by: Ali, Muhammad, et al.
Published: (2024)
by: Ali, Muhammad, et al.
Published: (2024)
Large Language Models are Strong Audio-Visual Speech Recognition Learners
by: Cappellazzo, Umberto, et al.
Published: (2024)
by: Cappellazzo, Umberto, et al.
Published: (2024)
ELGC-Net: Efficient Local-Global Context Aggregation for Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
POLY-SIM: Polyglot Speaker Identification with Missing Modality Grand Challenge 2026 Evaluation Plan
by: Moscati, Marta, et al.
Published: (2026)
by: Moscati, Marta, et al.
Published: (2026)
Towards SAR Automatic Target Recognition MultiCategory SAR Image Classification Based on Light Weight Vision Transformer
by: Zhao, Guibin, et al.
Published: (2024)
by: Zhao, Guibin, et al.
Published: (2024)
CDChat: A Large Multimodal Model for Remote Sensing Change Description
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
A Novel Deep Hybrid Framework with Ensemble-Based Feature Optimization for Robust Real-Time Human Activity Recognition
by: Ullah, Wasi, et al.
Published: (2025)
by: Ullah, Wasi, et al.
Published: (2025)
Rethinking Transformers Pre-training for Multi-Spectral Satellite Imagery
by: Noman, Mubashir, et al.
Published: (2024)
by: Noman, Mubashir, et al.
Published: (2024)
EgoAdapt: Enhancing Robustness in Egocentric Interactive Speaker Detection Under Missing Modalities
by: Qian, Xinyuan, et al.
Published: (2026)
by: Qian, Xinyuan, et al.
Published: (2026)
Robust Light-Weight Facial Affective Behavior Recognition with CLIP
by: Lin, Li, et al.
Published: (2024)
by: Lin, Li, et al.
Published: (2024)
Boosting Gesture Recognition with an Automatic Gesture Annotation Framework
by: Shen, Junxiao, et al.
Published: (2024)
by: Shen, Junxiao, et al.
Published: (2024)
EMF: Event Meta Formers for Event-based Real-time Traffic Object Detection
by: Khan, Muhammad Ahmed Ullah, et al.
Published: (2025)
by: Khan, Muhammad Ahmed Ullah, et al.
Published: (2025)
Real-time Traffic Object Detection for Autonomous Driving
by: Khan, Abdul Hannan, et al.
Published: (2024)
by: Khan, Abdul Hannan, et al.
Published: (2024)
Exploring Light-Weight Object Recognition for Real-Time Document Detection
by: Wojcik, Lucas, et al.
Published: (2025)
by: Wojcik, Lucas, et al.
Published: (2025)
The Components of Collaborative Joint Perception and Prediction -- A Conceptual Framework
by: Wan, Lei, et al.
Published: (2025)
by: Wan, Lei, et al.
Published: (2025)
An Evaluation of Large Pre-Trained Models for Gesture Recognition using Synthetic Videos
by: Reddy, Arun, et al.
Published: (2024)
by: Reddy, Arun, et al.
Published: (2024)
Sample-aware RandAugment: Search-free Automatic Data Augmentation for Effective Image Recognition
by: Xiao, Anqi, et al.
Published: (2025)
by: Xiao, Anqi, et al.
Published: (2025)
Enhancing LLM-based Autonomous Driving with Modular Traffic Light and Sign Recognition
by: Schmidt, Fabian, et al.
Published: (2025)
by: Schmidt, Fabian, et al.
Published: (2025)
MSRS: Training Multimodal Speech Recognition Models from Scratch with Sparse Mask Optimization
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
Full-Rank No More: Low-Rank Weight Training for Modern Speech Recognition Models
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
by: Fernandez-Lopez, Adriana, et al.
Published: (2024)
StretchySnake: Flexible SSM Training Unlocks Action Recognition Across Spatio-Temporal Scales
by: Siddiqui, Nyle, et al.
Published: (2025)
by: Siddiqui, Nyle, et al.
Published: (2025)
oTTC: Object Time-to-Contact for Motion Estimation in Autonomous Driving
by: Khan, Abdul Hannan, et al.
Published: (2024)
by: Khan, Abdul Hannan, et al.
Published: (2024)
Automatic Image Unfolding and Stitching Framework for Esophageal Lining Video Based on Density-Weighted Feature Matching
by: Li, Muyang, et al.
Published: (2024)
by: Li, Muyang, et al.
Published: (2024)
GTAutoAct: An Automatic Datasets Generation Framework Based on Game Engine Redevelopment for Action Recognition
by: Song, Xingyu, et al.
Published: (2024)
by: Song, Xingyu, et al.
Published: (2024)
GLip: A Global-Local Integrated Progressive Framework for Robust Visual Speech Recognition
by: Wang, Tianyue, et al.
Published: (2025)
by: Wang, Tianyue, et al.
Published: (2025)
Saliency-Aware Automatic Buddhas Statue Recognition
by: Qi, Yong, et al.
Published: (2024)
by: Qi, Yong, et al.
Published: (2024)
An Application-Agnostic Automatic Target Recognition System Using Vision Language Models
by: Palladino, Anthony, et al.
Published: (2024)
by: Palladino, Anthony, et al.
Published: (2024)
Automatic Labelling for Low-Light Pedestrian Detection
by: Bouzoulas, Dimitrios, et al.
Published: (2025)
by: Bouzoulas, Dimitrios, et al.
Published: (2025)
Dataset Distillation by Automatic Training Trajectories
by: Liu, Dai, et al.
Published: (2024)
by: Liu, Dai, et al.
Published: (2024)
Similar Items
-
Distillation-based Layer Dropping (DLD): Effective End-to-end Framework for Dynamic Speech Networks
by: Hannan, Abdul, et al.
Published: (2026) -
Input Conditioned Layer Dropping in Speech Foundation Models
by: Hannan, Abdul, et al.
Published: (2025) -
EoCD: Encoder only Remote Sensing Change Detection
by: Noman, Mubashir, et al.
Published: (2026) -
PAEFF: Precise Alignment and Enhanced Gated Feature Fusion for Face-Voice Association
by: Hannan, Abdul, et al.
Published: (2025) -
RFOP: Rethinking Fusion and Orthogonal Projection for Face-Voice Association
by: Hannan, Abdul, et al.
Published: (2025)