ResidualTransformer: Residual Low-Rank Learning with Weight-Sharing for Transformer Layers
Fuente:
arXiv
Salvato in:
| Autori principali: | Wang, Yiming, Li, Jinyu |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
di: Kögel, Fabian, et al.
Pubblicazione: (2023)
di: Kögel, Fabian, et al.
Pubblicazione: (2023)
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
di: Wang, Xiaofei, et al.
Pubblicazione: (2023)
Bayesian Low-Rank Factorization for Robust Model Adaptation
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
Property Neurons in Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024)
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
di: Zhu, Haina, et al.
Pubblicazione: (2025)
di: Zhu, Haina, et al.
Pubblicazione: (2025)
Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers
di: Wang, Juncheng, et al.
Pubblicazione: (2025)
di: Wang, Juncheng, et al.
Pubblicazione: (2025)
How Redundant Is the Transformer Stack in Speech Representation Models?
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
Cascaded Cross-Modal Transformer for Audio-Textual Classification
di: Ristea, Nicolae-Catalin, et al.
Pubblicazione: (2024)
di: Ristea, Nicolae-Catalin, et al.
Pubblicazione: (2024)
Token-Weighted RNN-T for Learning from Flawed Data
di: Keren, Gil, et al.
Pubblicazione: (2024)
di: Keren, Gil, et al.
Pubblicazione: (2024)
Transformer-based Model for ASR N-Best Rescoring and Rewriting
di: Kang, Iwen E., et al.
Pubblicazione: (2024)
di: Kang, Iwen E., et al.
Pubblicazione: (2024)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech
di: Battenberg, Eric, et al.
Pubblicazione: (2024)
di: Battenberg, Eric, et al.
Pubblicazione: (2024)
Dysarthria Normalization via Local Lie Group Transformations for Robust ASR
di: Osipov, Mikhail
Pubblicazione: (2025)
di: Osipov, Mikhail
Pubblicazione: (2025)
Advancing Airport Tower Command Recognition: Integrating Squeeze-and-Excitation and Broadcasted Residual Learning
di: Lin, Yuanxi, et al.
Pubblicazione: (2024)
di: Lin, Yuanxi, et al.
Pubblicazione: (2024)
Variable Bitrate Residual Vector Quantization for Audio Coding
di: Chae, Yunkee, et al.
Pubblicazione: (2024)
di: Chae, Yunkee, et al.
Pubblicazione: (2024)
XLS-R Deep Learning Model for Multilingual ASR on Low- Resource Languages: Indonesian, Javanese, and Sundanese
di: Arisaputra, Panji, et al.
Pubblicazione: (2024)
di: Arisaputra, Panji, et al.
Pubblicazione: (2024)
Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
di: Liu, Zhijun, et al.
Pubblicazione: (2024)
Auto-adaptive Resonance Equalization using Dilated Residual Networks
di: Grachten, Maarten, et al.
Pubblicazione: (2018)
di: Grachten, Maarten, et al.
Pubblicazione: (2018)
Evaluating the Effectiveness of Transformer Layers in Wav2Vec 2.0, XLS-R, and Whisper for Speaker Identification Tasks
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
Attentive Fusion: A Transformer-based Approach to Multimodal Hate Speech Detection
di: Mandal, Atanu, et al.
Pubblicazione: (2024)
di: Mandal, Atanu, et al.
Pubblicazione: (2024)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
di: Likhomanenko, Tatiana, et al.
Pubblicazione: (2025)
di: Likhomanenko, Tatiana, et al.
Pubblicazione: (2025)
DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation
di: Jia, Dongya, et al.
Pubblicazione: (2025)
di: Jia, Dongya, et al.
Pubblicazione: (2025)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
Rasa: Building Expressive Speech Synthesis Systems for Indian Languages in Low-resource Settings
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024)
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2024)
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
di: Amooie, Reihaneh, et al.
Pubblicazione: (2025)
di: Amooie, Reihaneh, et al.
Pubblicazione: (2025)
Enhancing Out-of-Vocabulary Performance of Indian TTS Systems for Practical Applications through Low-Effort Data Strategies
di: Anand, Srija, et al.
Pubblicazione: (2024)
di: Anand, Srija, et al.
Pubblicazione: (2024)
Phir Hera Fairy: An English Fairytaler is a Strong Faker of Fluent Speech in Low-Resource Indian Languages
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2025)
di: Varadhan, Praveen Srinivasa, et al.
Pubblicazione: (2025)
AdaPTwin: Low-Cost Adaptive Compression of Product Twins in Transformers
di: Biju, Emil, et al.
Pubblicazione: (2024)
di: Biju, Emil, et al.
Pubblicazione: (2024)
Siamese Residual Neural Network for Musical Shape Evaluation in Piano Performance Assessment
di: Li, Xiaoquan, et al.
Pubblicazione: (2024)
di: Li, Xiaoquan, et al.
Pubblicazione: (2024)
Federated Learning of Large ASR Models in the Real World
di: Xiao, Yonghui, et al.
Pubblicazione: (2024)
di: Xiao, Yonghui, et al.
Pubblicazione: (2024)
Meta Learning Text-to-Speech Synthesis in over 7000 Languages
di: Lux, Florian, et al.
Pubblicazione: (2024)
di: Lux, Florian, et al.
Pubblicazione: (2024)
EAT: Self-Supervised Pre-Training with Efficient Audio Transformer
di: Chen, Wenxi, et al.
Pubblicazione: (2024)
di: Chen, Wenxi, et al.
Pubblicazione: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
di: Li, Xingyuan, et al.
Pubblicazione: (2024)
di: Li, Xingyuan, et al.
Pubblicazione: (2024)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
di: Nagpal, Chirag, et al.
Pubblicazione: (2024)
ESC: Efficient Speech Coding with Cross-Scale Residual Vector Quantized Transformers
di: Gu, Yuzhe, et al.
Pubblicazione: (2024)
di: Gu, Yuzhe, et al.
Pubblicazione: (2024)
Learn and Don't Forget: Adding a New Language to ASR Foundation Models
di: Qian, Mengjie, et al.
Pubblicazione: (2024)
di: Qian, Mengjie, et al.
Pubblicazione: (2024)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
di: Zampierin, Luca, et al.
Pubblicazione: (2024)
di: Zampierin, Luca, et al.
Pubblicazione: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
di: Lu, Yen-Ju, et al.
Pubblicazione: (2024)
di: Lu, Yen-Ju, et al.
Pubblicazione: (2024)
Robust and Explainable Depression Identification from Speech Using Vowel-Based Ensemble Learning Approaches
di: Feng, Kexin, et al.
Pubblicazione: (2024)
di: Feng, Kexin, et al.
Pubblicazione: (2024)
SCRAPL: Scattering Transform with Random Paths for Machine Learning
di: Mitcheltree, Christopher, et al.
Pubblicazione: (2026)
di: Mitcheltree, Christopher, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Towards Robust FastSpeech 2 by Modelling Residual Multimodality
di: Kögel, Fabian, et al.
Pubblicazione: (2023) -
SpeechX: Neural Codec Language Model as a Versatile Speech Transformer
di: Wang, Xiaofei, et al.
Pubblicazione: (2023) -
Bayesian Low-Rank Factorization for Robust Model Adaptation
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025) -
Property Neurons in Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2024) -
MuQ: Self-Supervised Music Representation Learning with Mel Residual Vector Quantization
di: Zhu, Haina, et al.
Pubblicazione: (2025)