SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Pu, Watanabe, Shinji, Van hamme, Hugo |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
SSVD: Structured SVD for Parameter-Efficient Fine-Tuning and Benchmarking under Domain Shift in ASR
por: Wang, Pu, et al.
Publicado: (2025)
por: Wang, Pu, et al.
Publicado: (2025)
Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2022)
por: Eeckt, Steven Vander, et al.
Publicado: (2022)
Continual Learning for Monolingual End-to-End Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2021)
por: Eeckt, Steven Vander, et al.
Publicado: (2021)
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
por: Poncelet, Jakob, et al.
Publicado: (2025)
por: Poncelet, Jakob, et al.
Publicado: (2025)
Rehearsal-Free Online Continual Learning for Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2023)
por: Eeckt, Steven Vander, et al.
Publicado: (2023)
Generating Data with Text-to-Speech and Large-Language Models for Conversational Speech Recognition
por: Cornell, Samuele, et al.
Publicado: (2024)
por: Cornell, Samuele, et al.
Publicado: (2024)
Efficient Rehearsal for Continual Learning in ASR via Singular Value Tuning
por: Eeckt, Steven Vander, et al.
Publicado: (2026)
por: Eeckt, Steven Vander, et al.
Publicado: (2026)
OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification
por: Peng, Yifan, et al.
Publicado: (2024)
por: Peng, Yifan, et al.
Publicado: (2024)
Efficient Extraction of Noise-Robust Discrete Units from Self-Supervised Speech Models
por: Poncelet, Jakob, et al.
Publicado: (2024)
por: Poncelet, Jakob, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Dynamic Vocabulary
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
An Empirical Recipe for Universal Phone Recognition
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
por: Bharadwaj, Shikhar, et al.
Publicado: (2026)
Multi-blank Transducers for Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2022)
por: Xu, Hainan, et al.
Publicado: (2022)
Contextualized End-to-end Automatic Speech Recognition with Intermediate Biasing Loss
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
por: Wang, Qingzheng, et al.
Publicado: (2025)
por: Wang, Qingzheng, et al.
Publicado: (2025)
HyperTTS: Parameter Efficient Adaptation in Text to Speech using Hypernetworks
por: Li, Yingting, et al.
Publicado: (2024)
por: Li, Yingting, et al.
Publicado: (2024)
Comparison of Self-Supervised Speech Pre-Training Methods on Flemish Dutch
por: Poncelet, Jakob, et al.
Publicado: (2021)
por: Poncelet, Jakob, et al.
Publicado: (2021)
Universal Robust Speech Adaptation for Cross-Domain Speech Recognition and Enhancement
por: Wang, Chien-Chun, et al.
Publicado: (2026)
por: Wang, Chien-Chun, et al.
Publicado: (2026)
Joint Optimization of Streaming and Non-Streaming Automatic Speech Recognition with Multi-Decoder and Knowledge Distillation
por: Shakeel, Muhammad, et al.
Publicado: (2024)
por: Shakeel, Muhammad, et al.
Publicado: (2024)
Unsupervised Accent Adaptation Through Masked Language Model Correction Of Discrete Self-Supervised Speech Units
por: Poncelet, Jakob, et al.
Publicado: (2023)
por: Poncelet, Jakob, et al.
Publicado: (2023)
Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data
por: Choi, Youngwon, et al.
Publicado: (2025)
por: Choi, Youngwon, et al.
Publicado: (2025)
Hypothesis Clustering and Merging: Novel MultiTalker Speech Recognition with Speaker Tokens
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
por: Sudo, Yui, et al.
Publicado: (2024)
por: Sudo, Yui, et al.
Publicado: (2024)
Speech Robust Bench: A Robustness Benchmark For Speech Recognition
por: Shah, Muhammad A., et al.
Publicado: (2024)
por: Shah, Muhammad A., et al.
Publicado: (2024)
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
por: Inoue, Nakamasa, et al.
Publicado: (2024)
por: Inoue, Nakamasa, et al.
Publicado: (2024)
Spiralformer: Low Latency Encoder for Streaming Speech Recognition with Circular Layer Skipping and Early Exiting
por: Tsunoo, Emiru, et al.
Publicado: (2025)
por: Tsunoo, Emiru, et al.
Publicado: (2025)
Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
por: Kashiwagi, Yosuke, et al.
Publicado: (2024)
Moonshine: Speech Recognition for Live Transcription and Voice Commands
por: Jeffries, Nat, et al.
Publicado: (2024)
por: Jeffries, Nat, et al.
Publicado: (2024)
Speech Recognition With LLMs Adapted to Disordered Speech Using Reinforcement Learning
por: Nagpal, Chirag, et al.
Publicado: (2024)
por: Nagpal, Chirag, et al.
Publicado: (2024)
DYNAC: Dynamic Vocabulary based Non-Autoregressive Contextualization for Speech Recognition
por: Sudo, Yui, et al.
Publicado: (2025)
por: Sudo, Yui, et al.
Publicado: (2025)
On the Problem of Text-To-Speech Model Selection for Synthetic Data Generation in Automatic Speech Recognition
por: Rossenbach, Nick, et al.
Publicado: (2024)
por: Rossenbach, Nick, et al.
Publicado: (2024)
Pre-Finetuning for Few-Shot Emotional Speech Recognition
por: Chen, Maximillian, et al.
Publicado: (2023)
por: Chen, Maximillian, et al.
Publicado: (2023)
Transducers with Pronunciation-aware Embeddings for Automatic Speech Recognition
por: Xu, Hainan, et al.
Publicado: (2024)
por: Xu, Hainan, et al.
Publicado: (2024)
Regularizing Learnable Feature Extraction for Automatic Speech Recognition
por: Vieting, Peter, et al.
Publicado: (2025)
por: Vieting, Peter, et al.
Publicado: (2025)
Windowed SummaryMixing: An Efficient Fine-Tuning of Self-Supervised Learning Models for Low-resource Speech Recognition
por: Menon, Aditya Srinivas, et al.
Publicado: (2026)
por: Menon, Aditya Srinivas, et al.
Publicado: (2026)
Gradient Norm-based Fine-Tuning for Backdoor Defense in Automatic Speech Recognition
por: Zhou, Nanjun, et al.
Publicado: (2025)
por: Zhou, Nanjun, et al.
Publicado: (2025)
Parameter Efficient Finetuning for Speech Emotion Recognition and Domain Adaptation
por: Lashkarashvili, Nineli, et al.
Publicado: (2024)
por: Lashkarashvili, Nineli, et al.
Publicado: (2024)
Examining Test-Time Adaptation for Personalized Child Speech Recognition
por: Shi, Zhonghao, et al.
Publicado: (2024)
por: Shi, Zhonghao, et al.
Publicado: (2024)
Transcription-Free Fine-Tuning of Speech Separation Models for Noisy and Reverberant Multi-Speaker Automatic Speech Recognition
por: Ravenscroft, William, et al.
Publicado: (2024)
por: Ravenscroft, William, et al.
Publicado: (2024)
On-device Streaming Discrete Speech Units
por: Choi, Kwanghee, et al.
Publicado: (2025)
por: Choi, Kwanghee, et al.
Publicado: (2025)
OLMoASR: Open Models and Data for Training Robust Speech Recognition Models
por: Ngo, Huong, et al.
Publicado: (2025)
por: Ngo, Huong, et al.
Publicado: (2025)
Ejemplares similares
-
SSVD: Structured SVD for Parameter-Efficient Fine-Tuning and Benchmarking under Domain Shift in ASR
por: Wang, Pu, et al.
Publicado: (2025) -
Weight Averaging: A Simple Yet Effective Method to Overcome Catastrophic Forgetting in Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2022) -
Continual Learning for Monolingual End-to-End Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2021) -
Leveraging Broadcast Media Subtitle Transcripts for Automatic Speech Recognition and Subtitling
por: Poncelet, Jakob, et al.
Publicado: (2025) -
Rehearsal-Free Online Continual Learning for Automatic Speech Recognition
por: Eeckt, Steven Vander, et al.
Publicado: (2023)