Efficient Supernet Training with Orthogonal Softmax for Scalable ASR Model Compression
Fuente:
arXiv
Guardado en:
| Autores principales: | Xu, Jingjing, Beck, Eugen, Yang, Zijian, Schlüter, Ralf |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Dynamic Acoustic Model Architecture Optimization in Training for ASR
por: Xu, Jingjing, et al.
Publicado: (2025)
por: Xu, Jingjing, et al.
Publicado: (2025)
Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
por: Xu, Jingjing, et al.
Publicado: (2024)
por: Xu, Jingjing, et al.
Publicado: (2024)
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
por: Jawahar, Ganesh, et al.
Publicado: (2023)
por: Jawahar, Ganesh, et al.
Publicado: (2023)
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
por: Yang, Zijian, et al.
Publicado: (2023)
por: Yang, Zijian, et al.
Publicado: (2023)
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
por: Beck, Eugen, et al.
Publicado: (2026)
por: Beck, Eugen, et al.
Publicado: (2026)
Scalable-Softmax Is Superior for Attention
por: Nakanishi, Ken M.
Publicado: (2025)
por: Nakanishi, Ken M.
Publicado: (2025)
Unified Learnable 2D Convolutional Feature Extraction for ASR
por: Vieting, Peter, et al.
Publicado: (2025)
por: Vieting, Peter, et al.
Publicado: (2025)
Label-Context-Dependent Internal Language Model Estimation for CTC
por: Yang, Zijian, et al.
Publicado: (2025)
por: Yang, Zijian, et al.
Publicado: (2025)
Progressive Supernet Training for Efficient Visual Autoregressive Modeling
por: Chen, Xiaoyue, et al.
Publicado: (2025)
por: Chen, Xiaoyue, et al.
Publicado: (2025)
Multi-agent Architecture Search via Agentic Supernet
por: Zhang, Guibin, et al.
Publicado: (2025)
por: Zhang, Guibin, et al.
Publicado: (2025)
Streaming Bilingual End-to-End ASR model using Attention over Multiple Softmax
por: Patil, Aditya, et al.
Publicado: (2024)
por: Patil, Aditya, et al.
Publicado: (2024)
Self-Adjust Softmax
por: Zheng, Chuanyang, et al.
Publicado: (2025)
por: Zheng, Chuanyang, et al.
Publicado: (2025)
Scalable Vision Language Model Training via High Quality Data Curation
por: Dong, Hongyuan, et al.
Publicado: (2025)
por: Dong, Hongyuan, et al.
Publicado: (2025)
Compressible Softmax-Attended Language under Incompressible Attention
por: Lee, Wonsuk
Publicado: (2026)
por: Lee, Wonsuk
Publicado: (2026)
On the Effect of Purely Synthetic Training Data for Different Automatic Speech Recognition Architectures
por: Hilmes, Benedikt, et al.
Publicado: (2024)
por: Hilmes, Benedikt, et al.
Publicado: (2024)
UniAttn: Reducing Inference Costs via Softmax Unification for Post-Training LLMs
por: Xiong, Yizhe, et al.
Publicado: (2025)
por: Xiong, Yizhe, et al.
Publicado: (2025)
MSA-ASR: Efficient Multilingual Speaker Attribution with frozen ASR Models
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
por: Nguyen, Thai-Binh, et al.
Publicado: (2024)
To Softmax, or not to Softmax: that is the question when applying Active Learning for Transformer Models
por: Gonsior, Julius, et al.
Publicado: (2022)
por: Gonsior, Julius, et al.
Publicado: (2022)
Building Robust and Scalable Multilingual ASR for Indian Languages
por: Gangwar, Arjun, et al.
Publicado: (2025)
por: Gangwar, Arjun, et al.
Publicado: (2025)
TensorGPT: Efficient Compression of Large Language Models based on Tensor-Train Decomposition
por: Xu, Mingxue, et al.
Publicado: (2023)
por: Xu, Mingxue, et al.
Publicado: (2023)
Polyglot-Lion: Efficient Multilingual ASR for Singapore via Balanced Fine-Tuning of Qwen3-ASR
por: Dang, Quy-Anh, et al.
Publicado: (2026)
por: Dang, Quy-Anh, et al.
Publicado: (2026)
Orthogonal Finetuning Made Scalable
por: Qiu, Zeju, et al.
Publicado: (2025)
por: Qiu, Zeju, et al.
Publicado: (2025)
Efficiently Train ASR Models that Memorize Less and Perform Better with Per-core Clipping
por: Wang, Lun, et al.
Publicado: (2024)
por: Wang, Lun, et al.
Publicado: (2024)
Counting Like Transformers: Compiling Temporal Counting Logic Into Softmax Transformers
por: Yang, Andy, et al.
Publicado: (2024)
por: Yang, Andy, et al.
Publicado: (2024)
UORA: Uniform Orthogonal Reinitialization Adaptation in Parameter-Efficient Fine-Tuning of Large Models
por: Zhang, Xueyan, et al.
Publicado: (2025)
por: Zhang, Xueyan, et al.
Publicado: (2025)
Scalable Offline ASR for Command-Style Dictation in Courtrooms
por: Nethil, Kumarmanas, et al.
Publicado: (2025)
por: Nethil, Kumarmanas, et al.
Publicado: (2025)
Exploring Gender Bias in Large Language Models: An In-depth Dive into the German Language
por: Gnadt, Kristin, et al.
Publicado: (2025)
por: Gnadt, Kristin, et al.
Publicado: (2025)
Supplementary Resources and Analysis for Automatic Speech Recognition Systems Trained on the Loquacious Dataset
por: Rossenbach, Nick, et al.
Publicado: (2025)
por: Rossenbach, Nick, et al.
Publicado: (2025)
X-EcoMLA: Upcycling Pre-Trained Attention into MLA for Efficient and Extreme KV Compression
por: Li, Guihong, et al.
Publicado: (2025)
por: Li, Guihong, et al.
Publicado: (2025)
Text-Utilization for Encoder-dominated Speech Recognition Models
por: Zeyer, Albert, et al.
Publicado: (2026)
por: Zeyer, Albert, et al.
Publicado: (2026)
BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding
por: Yuan, Jiayi, et al.
Publicado: (2025)
por: Yuan, Jiayi, et al.
Publicado: (2025)
Functional Abstraction of Knowledge Recall in Large Language Models
por: Wang, Zijian, et al.
Publicado: (2025)
por: Wang, Zijian, et al.
Publicado: (2025)
ProcrustesGPT: Compressing LLMs with Structured Matrices and Orthogonal Transformations
por: Grishina, Ekaterina, et al.
Publicado: (2025)
por: Grishina, Ekaterina, et al.
Publicado: (2025)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
por: Rossenbach, Nick, et al.
Publicado: (2024)
por: Rossenbach, Nick, et al.
Publicado: (2024)
Efficient Adaptation of Multilingual Models for Japanese ASR
por: Bajo, Mark, et al.
Publicado: (2024)
por: Bajo, Mark, et al.
Publicado: (2024)
Diffusion Language Models for Speech Recognition
por: Naveriani, Davyd, et al.
Publicado: (2026)
por: Naveriani, Davyd, et al.
Publicado: (2026)
PromptASR for contextualized ASR with controllable style
por: Yang, Xiaoyu, et al.
Publicado: (2023)
por: Yang, Xiaoyu, et al.
Publicado: (2023)
Parameter Efficient Quasi-Orthogonal Fine-Tuning via Givens Rotation
por: Ma, Xinyu, et al.
Publicado: (2024)
por: Ma, Xinyu, et al.
Publicado: (2024)
Universal Adversarial Suffixes Using Calibrated Gumbel-Softmax Relaxation
por: Soor, Sampriti, et al.
Publicado: (2025)
por: Soor, Sampriti, et al.
Publicado: (2025)
Forgetting Transformer: Softmax Attention with a Forget Gate
por: Lin, Zhixuan, et al.
Publicado: (2025)
por: Lin, Zhixuan, et al.
Publicado: (2025)
Ejemplares similares
-
Dynamic Acoustic Model Architecture Optimization in Training for ASR
por: Xu, Jingjing, et al.
Publicado: (2025) -
Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition
por: Xu, Jingjing, et al.
Publicado: (2024) -
Mixture-of-Supernets: Improving Weight-Sharing Supernet Training with Architecture-Routed Mixture-of-Experts
por: Jawahar, Ganesh, et al.
Publicado: (2023) -
On the Relation between Internal Language Model and Sequence Discriminative Training for Neural Transducers
por: Yang, Zijian, et al.
Publicado: (2023) -
AppTek Call-Center Dialogues: A Multi-Accent Long-Form Benchmark for English ASR
por: Beck, Eugen, et al.
Publicado: (2026)