StableQuant: Layer Adaptive Post-Training Quantization for Speech Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Yeona, Han, Hyewon, Chung, Woo-jin, Kang, Hong-Goo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Optimization of DNN-based speaker verification model through efficient quantization technique
von: Hong, Yeona, et al.
Veröffentlicht: (2024)
von: Hong, Yeona, et al.
Veröffentlicht: (2024)
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
von: Han, Wooseok, et al.
Veröffentlicht: (2024)
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
von: Chung, Woo-Jin, et al.
Veröffentlicht: (2023)
von: Chung, Woo-Jin, et al.
Veröffentlicht: (2023)
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
Speaker-Independent Acoustic-to-Articulatory Inversion through Multi-Channel Attention Discriminator
von: Chung, Woo-Jin, et al.
Veröffentlicht: (2024)
von: Chung, Woo-Jin, et al.
Veröffentlicht: (2024)
Unifying Model and Layer Fusion for Speech Foundation Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2025)
Neural Spectral Band Generation for Audio Coding
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)
Layer-aware TDNN: Speaker Recognition Using Multi-Layer Features from Pre-Trained Models
von: Kim, Jin Sob, et al.
Veröffentlicht: (2024)
von: Kim, Jin Sob, et al.
Veröffentlicht: (2024)
Text-To-Speech Synthesis In The Wild
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
One-pass Multiple Conformer and Foundation Speech Systems Compression and Quantization Using An All-in-one Neural Model
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
von: Li, Zhaoqing, et al.
Veröffentlicht: (2024)
LAMA-UT: Language Agnostic Multilingual ASR through Orthography Unification and Language-Specific Transliteration
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
Enhanced Deep Speech Separation in Clustered Ad Hoc Distributed Microphone Environments
von: Kim, Jihyun, et al.
Veröffentlicht: (2024)
von: Kim, Jihyun, et al.
Veröffentlicht: (2024)
PS-TTS: Phonetic Synchronization in Text-to-Speech for Achieving Natural Automated Dubbing
von: Hong, Changi, et al.
Veröffentlicht: (2026)
von: Hong, Changi, et al.
Veröffentlicht: (2026)
Zero-Shot Text-to-Speech as Golden Speech Generator: A Systematic Framework and its Applicability in Automatic Pronunciation Assessment
von: Lo, Tien-Hong, et al.
Veröffentlicht: (2024)
von: Lo, Tien-Hong, et al.
Veröffentlicht: (2024)
UniCoM: A Universal Code-Switching Speech Generator
von: Lee, Sangmin, et al.
Veröffentlicht: (2025)
von: Lee, Sangmin, et al.
Veröffentlicht: (2025)
Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
von: Kim, Miseul, et al.
Veröffentlicht: (2024)
von: Kim, Miseul, et al.
Veröffentlicht: (2024)
Adaptive Duration Model for Text Speech Alignment
von: Cao, Junjie
Veröffentlicht: (2025)
von: Cao, Junjie
Veröffentlicht: (2025)
SpoofCeleb: Speech Deepfake Detection and SASV In The Wild
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
Unveiling the Best Practices for Applying Speech Foundation Models to Speech Intelligibility Prediction for Hearing-Impaired People
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
von: Zhou, Haoshuai, et al.
Veröffentlicht: (2025)
Speech-to-Speech Translation with Discrete-Unit-Based Style Transfer
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
von: Wang, Yongqi, et al.
Veröffentlicht: (2023)
CrossSpeech++: Cross-lingual Speech Synthesis with Decoupled Language and Speaker Generation
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
von: Kim, Ji-Hoon, et al.
Veröffentlicht: (2024)
To what extent can ASV systems naturally defend against spoofing attacks?
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
DEX-TTS: Diffusion-based EXpressive Text-to-Speech with Style Modeling on Time Variability
von: Park, Hyun Joon, et al.
Veröffentlicht: (2024)
von: Park, Hyun Joon, et al.
Veröffentlicht: (2024)
Improvement and Implementation of a Speech Emotion Recognition Model Based on Dual-Layer LSTM
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
von: Yang, Xiaoran, et al.
Veröffentlicht: (2024)
Evaluating the Effectiveness of Pre-Trained Audio Embeddings for Classification of Parkinson's Disease Speech Data
von: Postma, Emmy, et al.
Veröffentlicht: (2025)
von: Postma, Emmy, et al.
Veröffentlicht: (2025)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Foundation Model Training on EU Languages
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
von: Gaido, Marco, et al.
Veröffentlicht: (2024)
HARNESS: Lightweight Distilled Arabic Speech Foundation Models
von: Sukhadia, Vrunda N., et al.
Veröffentlicht: (2026)
von: Sukhadia, Vrunda N., et al.
Veröffentlicht: (2026)
MERaLiON-SpeechEncoder: Towards a Speech Foundation Model for Singapore and Beyond
von: Huzaifah, Muhammad, et al.
Veröffentlicht: (2024)
von: Huzaifah, Muhammad, et al.
Veröffentlicht: (2024)
Thinking in Directivity: Speech Large Language Model for Multi-Talker Directional Speech Recognition
von: Xie, Jiamin, et al.
Veröffentlicht: (2025)
von: Xie, Jiamin, et al.
Veröffentlicht: (2025)
Noise-Agnostic Multitask Whisper Training for Reducing False Alarm Errors in Call-for-Help Detection
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2025)
von: Ryu, Myeonghoon, et al.
Veröffentlicht: (2025)
Speech Foundation Model Ensembles for the Controlled Singing Voice Deepfake Detection (CtrSVDD) Challenge 2024
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)
von: Guragain, Anmol, et al.
Veröffentlicht: (2024)
Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-spoofing
von: Liu, Tianchi, et al.
Veröffentlicht: (2025)
von: Liu, Tianchi, et al.
Veröffentlicht: (2025)
Speech Recognition on TV Series with Video-guided Post-ASR Correction
von: Yang, Haoyuan, et al.
Veröffentlicht: (2025)
von: Yang, Haoyuan, et al.
Veröffentlicht: (2025)
Gesture-Aware Zero-Shot Speech Recognition for Patients with Language Disorders
von: Kim, Seungbae, et al.
Veröffentlicht: (2025)
von: Kim, Seungbae, et al.
Veröffentlicht: (2025)
ASVspoof 5: Crowdsourced Speech Data, Deepfakes, and Adversarial Attacks at Scale
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Effective and Efficient One-pass Compression of Speech Foundation Models Using Sparsity-aware Self-pinching Gates
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
von: Xu, Haoning, et al.
Veröffentlicht: (2025)
Adaptive Knowledge Distillation for Device-Directed Speech Detection
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
Investigating Stochastic Methods for Prosody Modeling in Speech Synthesis
von: Mayer, Paul, et al.
Veröffentlicht: (2025)
von: Mayer, Paul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Optimization of DNN-based speaker verification model through efficient quantization technique
von: Hong, Yeona, et al.
Veröffentlicht: (2024) -
Stable-TTS: Stable Speaker-Adaptive Text-to-Speech Synthesis via Prosody Prompting
von: Han, Wooseok, et al.
Veröffentlicht: (2024) -
Effective and Efficient Mixed Precision Quantization of Speech Foundation Models
von: Xu, Haoning, et al.
Veröffentlicht: (2025) -
MF-PAM: Accurate Pitch Estimation through Periodicity Analysis and Multi-level Feature Fusion
von: Chung, Woo-Jin, et al.
Veröffentlicht: (2023) -
UniverSR: Unified and Versatile Audio Super-Resolution via Vocoder-Free Flow Matching
von: Choi, Woongjib, et al.
Veröffentlicht: (2025)