Scaling Properties of Speech Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cuervo, Santiago, Marxer, Ricard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024)
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024)
Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition
von: Yu, Yu, et al.
Veröffentlicht: (2024)
von: Yu, Yu, et al.
Veröffentlicht: (2024)
Low-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition
von: Yu, Yu, et al.
Veröffentlicht: (2023)
von: Yu, Yu, et al.
Veröffentlicht: (2023)
How to Estimate Model Transferability of Pre-Trained Speech Models?
von: Chen, Zih-Ching, et al.
Veröffentlicht: (2023)
von: Chen, Zih-Ching, et al.
Veröffentlicht: (2023)
Automatic Voice Identification after Speech Resynthesis using PPG
von: Gaudier, Thibault, et al.
Veröffentlicht: (2024)
von: Gaudier, Thibault, et al.
Veröffentlicht: (2024)
Text Injection for Neural Contextual Biasing
von: Meng, Zhong, et al.
Veröffentlicht: (2024)
von: Meng, Zhong, et al.
Veröffentlicht: (2024)
Deferred NAM: Low-latency Top-K Context Injection via Deferred Context Encoding for Non-Streaming ASR
von: Wu, Zelin, et al.
Veröffentlicht: (2024)
von: Wu, Zelin, et al.
Veröffentlicht: (2024)
Late Fusion and Multi-Level Fission Amplify Cross-Modal Transfer in Text-Speech LMs
von: Cuervo, Santiago, et al.
Veröffentlicht: (2025)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2025)
Global-Local Convolution with Spiking Neural Networks for Energy-efficient Keyword Spotting
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Robust online reconstruction of continuous-time signals from a lean spike train ensemble code
von: Chattopadhyay, Anik, et al.
Veröffentlicht: (2024)
von: Chattopadhyay, Anik, et al.
Veröffentlicht: (2024)
Parsing Musical Structure to Enable Meaningful Variations
von: Kanani, Maziar, et al.
Veröffentlicht: (2025)
von: Kanani, Maziar, et al.
Veröffentlicht: (2025)
Transfer Learning from Whisper for Microscopic Intelligibility Prediction
von: Best, Paul, et al.
Veröffentlicht: (2024)
von: Best, Paul, et al.
Veröffentlicht: (2024)
DeepSpeech models show Human-like Performance and Processing of Cochlear Implant Inputs
von: Steinhardt, Cynthia R., et al.
Veröffentlicht: (2024)
von: Steinhardt, Cynthia R., et al.
Veröffentlicht: (2024)
Resource-Efficient Speech Quality Prediction through Quantization Aware Training and Binary Activation Maps
von: Nilsson, Mattias, et al.
Veröffentlicht: (2024)
von: Nilsson, Mattias, et al.
Veröffentlicht: (2024)
A novel Reservoir Architecture for Periodic Time Series Prediction
von: Yuan, Zhongju, et al.
Veröffentlicht: (2024)
von: Yuan, Zhongju, et al.
Veröffentlicht: (2024)
Long-Form Text-to-Music Generation with Adaptive Prompts: A Case Study in Tabletop Role-Playing Games Soundtracks
von: Marra, Felipe, et al.
Veröffentlicht: (2024)
von: Marra, Felipe, et al.
Veröffentlicht: (2024)
Learning Delays in Spiking Neural Networks using Dilated Convolutions with Learnable Spacings
von: Hammouamri, Ilyass, et al.
Veröffentlicht: (2023)
von: Hammouamri, Ilyass, et al.
Veröffentlicht: (2023)
Spoken Conversational Agents with Large Language Models
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
von: Yang, Chao-Han Huck, et al.
Veröffentlicht: (2025)
Accurate Mapping of RNNs on Neuromorphic Hardware with Adaptive Spiking Neurons
von: Boeshertz, Gauthier, et al.
Veröffentlicht: (2024)
von: Boeshertz, Gauthier, et al.
Veröffentlicht: (2024)
A Comparison of Temporal Encoders for Neuromorphic Keyword Spotting with Few Neurons
von: Nilsson, Mattias, et al.
Veröffentlicht: (2023)
von: Nilsson, Mattias, et al.
Veröffentlicht: (2023)
LMUFormer: Low Complexity Yet Powerful Spiking Model With Legendre Memory Units
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
von: Liu, Zeyu, et al.
Veröffentlicht: (2024)
Deep Photonic Reservoir Computer for Speech Recognition
von: Picco, Enrico, et al.
Veröffentlicht: (2023)
von: Picco, Enrico, et al.
Veröffentlicht: (2023)
Artificial Neural Networks Trained on Noisy Speech Exhibit the McGurk Effect
von: Grasse, Lukas, et al.
Veröffentlicht: (2024)
von: Grasse, Lukas, et al.
Veröffentlicht: (2024)
HyperSound: Generating Implicit Neural Representations of Audio Signals with Hypernetworks
von: Szatkowski, Filip, et al.
Veröffentlicht: (2022)
von: Szatkowski, Filip, et al.
Veröffentlicht: (2022)
Neurobench: DCASE 2020 Acoustic Scene Classification benchmark on XyloAudio 2
von: Ke, Weijie, et al.
Veröffentlicht: (2024)
von: Ke, Weijie, et al.
Veröffentlicht: (2024)
Low-power SNN-based audio source localisation using a Hilbert Transform spike encoding scheme
von: Haghighatshoar, Saeid, et al.
Veröffentlicht: (2024)
von: Haghighatshoar, Saeid, et al.
Veröffentlicht: (2024)
sVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks
von: Yang, Qu, et al.
Veröffentlicht: (2024)
von: Yang, Qu, et al.
Veröffentlicht: (2024)
Grammatical Structure and Grammatical Variations in Non-Metric Iranian Classical Music
von: Kanani, Maziar, et al.
Veröffentlicht: (2025)
von: Kanani, Maziar, et al.
Veröffentlicht: (2025)
Generative Voice Bursts during Phone Call
von: Ranjan, Paritosh, et al.
Veröffentlicht: (2025)
von: Ranjan, Paritosh, et al.
Veröffentlicht: (2025)
Spiking-LEAF: A Learnable Auditory front-end for Spiking Neural Networks
von: Song, Zeyang, et al.
Veröffentlicht: (2023)
von: Song, Zeyang, et al.
Veröffentlicht: (2023)
A Novel Transfer Learning Approach for Mental Stability Classification from Voice Signal
von: Islam, Rafiul, et al.
Veröffentlicht: (2026)
von: Islam, Rafiul, et al.
Veröffentlicht: (2026)
DPSNN: Spiking Neural Network for Low-Latency Streaming Speech Enhancement
von: Sun, Tao, et al.
Veröffentlicht: (2024)
von: Sun, Tao, et al.
Veröffentlicht: (2024)
Learning spatial hearing via innate mechanisms
von: Chu, Yang, et al.
Veröffentlicht: (2020)
von: Chu, Yang, et al.
Veröffentlicht: (2020)
Ternary Spike-based Neuromorphic Signal Processing System
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
von: Wang, Shuai, et al.
Veröffentlicht: (2024)
Spiketrum: An FPGA-based Implementation of a Neuromorphic Cochlea
von: Alsakkal, MHD Anas, et al.
Veröffentlicht: (2024)
von: Alsakkal, MHD Anas, et al.
Veröffentlicht: (2024)
Delayed Memory Unit: Modelling Temporal Dependency Through Delay Gate
von: Sun, Pengfei, et al.
Veröffentlicht: (2023)
von: Sun, Pengfei, et al.
Veröffentlicht: (2023)
Speech foundation models on intelligibility prediction for hearing-impaired listeners
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2024)
Closing the Gap Between Text and Speech Understanding in LLMs
von: Cuervo, Santiago, et al.
Veröffentlicht: (2025)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2025)
Parallel Stacked Aggregated Network for Voice Authentication in IoT-Enabled Smart Devices
von: Khan, Awais, et al.
Veröffentlicht: (2024)
von: Khan, Awais, et al.
Veröffentlicht: (2024)
Biomimetic Frontend for Differentiable Audio Processing
von: Famularo, Ruolan Leslie, et al.
Veröffentlicht: (2024)
von: Famularo, Ruolan Leslie, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hierarchical Recurrent Adapters for Efficient Multi-Task Adaptation of Large Speech Models
von: Munkhdalai, Tsendsuren, et al.
Veröffentlicht: (2024) -
Investigating Training Strategies and Model Robustness of Low-Rank Adaptation for Language Modeling in Speech Recognition
von: Yu, Yu, et al.
Veröffentlicht: (2024) -
Low-rank Adaptation of Large Language Model Rescoring for Parameter-Efficient Speech Recognition
von: Yu, Yu, et al.
Veröffentlicht: (2023) -
How to Estimate Model Transferability of Pre-Trained Speech Models?
von: Chen, Zih-Ching, et al.
Veröffentlicht: (2023) -
Automatic Voice Identification after Speech Resynthesis using PPG
von: Gaudier, Thibault, et al.
Veröffentlicht: (2024)