How to Estimate Model Transferability of Pre-Trained Speech Models?
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Zih-Ching, Yang, Chao-Han Huck, Li, Bo, Zhang, Yu, Chen, Nanxin, Chang, Shuo-Yiin, Prabhavalkar, Rohit, Lee, Hung-yi, Sainath, Tara N. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Text Injection for Neural Contextual Biasing
por: Meng, Zhong, et al.
Publicado: (2024)
por: Meng, Zhong, et al.
Publicado: (2024)
Resource-Efficient Speech Quality Prediction through Quantization Aware Training and Binary Activation Maps
por: Nilsson, Mattias, et al.
Publicado: (2024)
por: Nilsson, Mattias, et al.
Publicado: (2024)
Artificial Neural Networks Trained on Noisy Speech Exhibit the McGurk Effect
por: Grasse, Lukas, et al.
Publicado: (2024)
por: Grasse, Lukas, et al.
Publicado: (2024)
DeepSpeech models show Human-like Performance and Processing of Cochlear Implant Inputs
por: Steinhardt, Cynthia R., et al.
Publicado: (2024)
por: Steinhardt, Cynthia R., et al.
Publicado: (2024)
A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal
por: Islam, Rafiul, et al.
Publicado: (2026)
por: Islam, Rafiul, et al.
Publicado: (2026)
Spoken Conversational Agents with Large Language Models
por: Yang, Chao-Han Huck, et al.
Publicado: (2025)
por: Yang, Chao-Han Huck, et al.
Publicado: (2025)
Deep Photonic Reservoir Computer for Speech Recognition
por: Picco, Enrico, et al.
Publicado: (2023)
por: Picco, Enrico, et al.
Publicado: (2023)
Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition
por: Yu, Yu, et al.
Publicado: (2024)
por: Yu, Yu, et al.
Publicado: (2024)
Grammatical Structure and Grammatical Variations in Non-Metric Iranian Classical Music
por: Kanani, Maziar, et al.
Publicado: (2025)
por: Kanani, Maziar, et al.
Publicado: (2025)
Generative Voice Bursts during Phone Call
por: Ranjan, Paritosh, et al.
Publicado: (2025)
por: Ranjan, Paritosh, et al.
Publicado: (2025)
Spiking-LEAF: A Learnable Auditory front-end for Spiking Neural Networks
por: Song, Zeyang, et al.
Publicado: (2023)
por: Song, Zeyang, et al.
Publicado: (2023)
Neurobench: DCASE 2020 Acoustic Scene Classification benchmark on XyloAudio 2
por: Ke, Weijie, et al.
Publicado: (2024)
por: Ke, Weijie, et al.
Publicado: (2024)
Low-power SNN-based audio source localisation using a Hilbert Transform spike encoding scheme
por: Haghighatshoar, Saeid, et al.
Publicado: (2024)
por: Haghighatshoar, Saeid, et al.
Publicado: (2024)
sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
por: Yang, Qu, et al.
Publicado: (2024)
por: Yang, Qu, et al.
Publicado: (2024)
DPSNN: Spiking Neural Network for Low-Latency Streaming Speech Enhancement
por: Sun, Tao, et al.
Publicado: (2024)
por: Sun, Tao, et al.
Publicado: (2024)
Low-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition
por: Yu, Yu, et al.
Publicado: (2023)
por: Yu, Yu, et al.
Publicado: (2023)
Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models
por: Munkhdalai, Tsendsuren, et al.
Publicado: (2024)
por: Munkhdalai, Tsendsuren, et al.
Publicado: (2024)
Ternary Spike-based Neuromorphic Signal Processing System
por: Wang, Shuai, et al.
Publicado: (2024)
por: Wang, Shuai, et al.
Publicado: (2024)
Parsing Musical Structure to Enable Meaningful Variations
por: Kanani, Maziar, et al.
Publicado: (2025)
por: Kanani, Maziar, et al.
Publicado: (2025)
Parallel Stacked Aggregated Network for Voice Authentication in IoT-Enabled Smart Devices
por: Khan, Awais, et al.
Publicado: (2024)
por: Khan, Awais, et al.
Publicado: (2024)
Biomimetic Frontend for Differentiable Audio Processing
por: Famularo, Ruolan Leslie, et al.
Publicado: (2024)
por: Famularo, Ruolan Leslie, et al.
Publicado: (2024)
Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
por: Wang, Shuai, et al.
Publicado: (2024)
por: Wang, Shuai, et al.
Publicado: (2024)
Spiketrum: An FPGA-based Implementation of a Neuromorphic Cochlea
por: Alsakkal, MHD Anas, et al.
Publicado: (2024)
por: Alsakkal, MHD Anas, et al.
Publicado: (2024)
Robust online reconstruction of continuous-time signals from a lean spike train ensemble code
por: Chattopadhyay, Anik, et al.
Publicado: (2024)
por: Chattopadhyay, Anik, et al.
Publicado: (2024)
LVNS-RAVE: Diversified audio generation with RAVE and Latent Vector Novelty Search
por: Guo, Jinyue, et al.
Publicado: (2024)
por: Guo, Jinyue, et al.
Publicado: (2024)
LACTOSE: Linear Array of Conditions, TOpologies with Separated Error-backpropagation -- The Differentiable "IF" Conditional for Differentiable Digital Signal Processing
por: Clarke, Christopher Johann
Publicado: (2025)
por: Clarke, Christopher Johann
Publicado: (2025)
Spiking Music: Audio Compression with Event Based Auto-encoders
por: Lisboa, Martim, et al.
Publicado: (2024)
por: Lisboa, Martim, et al.
Publicado: (2024)
Automatic Voice Identification after Speech Resynthesis using PPG
por: Gaudier, Thibault, et al.
Publicado: (2024)
por: Gaudier, Thibault, et al.
Publicado: (2024)
LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory Units
por: Liu, Zeyu, et al.
Publicado: (2024)
por: Liu, Zeyu, et al.
Publicado: (2024)
Deformable Audio Transformer for Audio Event Detection
por: Zhu, Wentao
Publicado: (2023)
por: Zhu, Wentao
Publicado: (2023)
Long-Form Text-to-Music Generation with Adaptive Prompts: A Case Study in Tabletop Role-Playing Games Soundtracks
por: Marra, Felipe, et al.
Publicado: (2024)
por: Marra, Felipe, et al.
Publicado: (2024)
PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
por: Lin, Tzu-Han, et al.
Publicado: (2024)
por: Lin, Tzu-Han, et al.
Publicado: (2024)
Acoustic neural networks: Identifying design principles and exploring physical feasibility
por: Kalthoff, Ivan, et al.
Publicado: (2025)
por: Kalthoff, Ivan, et al.
Publicado: (2025)
Emotion Detection Using Conditional Generative Adversarial Networks (cGAN): A Deep Learning Approach
por: Srivastava, Anushka
Publicado: (2025)
por: Srivastava, Anushka
Publicado: (2025)
HyperSound: Generating Implicit Neural Representations of Audio Signals with Hypernetworks
por: Szatkowski, Filip, et al.
Publicado: (2022)
por: Szatkowski, Filip, et al.
Publicado: (2022)
Multilingual and Fully Non-Autoregressive ASR with Large Language Model Fusion: A Comprehensive Study
por: Huang, W. Ronny, et al.
Publicado: (2024)
por: Huang, W. Ronny, et al.
Publicado: (2024)
Scaling Properties of Speech Language Models
por: Cuervo, Santiago, et al.
Publicado: (2024)
por: Cuervo, Santiago, et al.
Publicado: (2024)
Extreme Encoder Output Frame Rate Reduction: Improving Computational Latencies of Large End-to-End Models
por: Prabhavalkar, Rohit, et al.
Publicado: (2024)
por: Prabhavalkar, Rohit, et al.
Publicado: (2024)
Dilated Convolution with Learnable Spacings
por: Khalfaoui-Hassani, Ismail
Publicado: (2024)
por: Khalfaoui-Hassani, Ismail
Publicado: (2024)
Accurate Mapping of RNNs on Neuromorphic Hardware with Adaptive Spiking Neurons
por: Boeshertz, Gauthier, et al.
Publicado: (2024)
por: Boeshertz, Gauthier, et al.
Publicado: (2024)
Ejemplares similares
-
Text Injection for Neural Contextual Biasing
por: Meng, Zhong, et al.
Publicado: (2024) -
Resource-Efficient Speech Quality Prediction through Quantization Aware Training and Binary Activation Maps
por: Nilsson, Mattias, et al.
Publicado: (2024) -
Artificial Neural Networks Trained on Noisy Speech Exhibit the McGurk Effect
por: Grasse, Lukas, et al.
Publicado: (2024) -
DeepSpeech models show Human-like Performance and Processing of Cochlear Implant Inputs
por: Steinhardt, Cynthia R., et al.
Publicado: (2024) -
A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal
por: Islam, Rafiul, et al.
Publicado: (2026)