STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jang, Kangwook, Kim, Sungnyun, Kim, Hoirin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
von: Kim, Sungnyun, et al.
Veröffentlicht: (2024)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2024)
One-Class Learning with Adaptive Centroid Shift for Audio Deepfake Detection
von: Kim, Hyun Myung, et al.
Veröffentlicht: (2024)
von: Kim, Hyun Myung, et al.
Veröffentlicht: (2024)
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
von: Wang, Yujin, et al.
Veröffentlicht: (2022)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
von: Hwang, Min-Jae, et al.
Veröffentlicht: (2024)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Lightweight Audio Segmentation for Long-form Speech Translation
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
von: Lee, Jaesong, et al.
Veröffentlicht: (2024)
Scaling Self-Supervised Speech Models Uncovers Deep Linguistic Relationships: Evidence from the Pacific Cluster
von: Kim, Minu, et al.
Veröffentlicht: (2026)
von: Kim, Minu, et al.
Veröffentlicht: (2026)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
von: Peng, Junyi, et al.
Veröffentlicht: (2025)
Interface Design for Self-Supervised Speech Models
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
von: Shih, Yi-Jen, et al.
Veröffentlicht: (2024)
SpeechGLUE: How Well Can Self-Supervised Speech Models Capture Linguistic Knowledge?
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2023)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
DiscreteSLU: A Large Language Model with Self-Supervised Discrete Speech Units for Spoken Language Understanding
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
von: Shon, Suwon, et al.
Veröffentlicht: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
von: Fan, Ruchao, et al.
Veröffentlicht: (2024)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
von: Kong, Jungil, et al.
Veröffentlicht: (2023)
Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models
von: Phuong, Tuan Dat, et al.
Veröffentlicht: (2025)
von: Phuong, Tuan Dat, et al.
Veröffentlicht: (2025)
Is Self-Supervised Learning Enough to Fill in the Gap? A Study on Speech Inpainting
von: Asaad, Ihab, et al.
Veröffentlicht: (2024)
von: Asaad, Ihab, et al.
Veröffentlicht: (2024)
How Far Do SSL Speech Models Listen for Tone? Temporal Focus of Tone Representation under Low-resource Transfer
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
MoHAVE: Mixture of Hierarchical Audio-Visual Experts for Robust Speech Recognition
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
Dub-S2ST: Textless Speech-to-Speech Translation for Seamless Dubbing
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
von: Choi, Jeongsoo, et al.
Veröffentlicht: (2025)
Efficient Interleaved Speech Modeling through Knowledge Distillation
von: Nouriborji, Mohammadmahdi, et al.
Veröffentlicht: (2025)
von: Nouriborji, Mohammadmahdi, et al.
Veröffentlicht: (2025)
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model
von: Park, Joonyong, et al.
Veröffentlicht: (2024)
von: Park, Joonyong, et al.
Veröffentlicht: (2024)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
von: Osakuade, Opeyemi, et al.
Veröffentlicht: (2024)
Multilingual Zero Resource Speech Recognition Base on Self-Supervise Pre-Trained Acoustic Models
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
von: Wang, Haoyu, et al.
Veröffentlicht: (2022)
Efficient Speech Translation through Model Compression and Knowledge Distillation
von: Moslem, Yasmin
Veröffentlicht: (2025)
von: Moslem, Yasmin
Veröffentlicht: (2025)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2025)
BiRQ: Bi-Level Self-Labeling Random Quantization for Self-Supervised Speech Recognition
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
von: Jiang, Liuyuan, et al.
Veröffentlicht: (2025)
What Do Speech Foundation Models Not Learn About Speech?
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
von: Waheed, Abdul, et al.
Veröffentlicht: (2024)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
von: Ashihara, Takanori, et al.
Veröffentlicht: (2024)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
von: Meghanani, Amit, et al.
Veröffentlicht: (2026)
Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
von: Xu, Tianyi, et al.
Veröffentlicht: (2025)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
von: Park, Chanho, et al.
Veröffentlicht: (2023)
von: Park, Chanho, et al.
Veröffentlicht: (2023)
Scalable Frameworks for Real-World Audio-Visual Speech Recognition
von: Kim, Sungnyun
Veröffentlicht: (2025)
von: Kim, Sungnyun
Veröffentlicht: (2025)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2025)
von: Venkateswaran, Nitin, et al.
Veröffentlicht: (2025)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
von: Hernandez, Abner, et al.
Veröffentlicht: (2026)
Paralinguistics-Aware Speech-Empowered Large Language Models for Natural Conversation
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
von: Kim, Heeseung, et al.
Veröffentlicht: (2024)
Re-Parameterization of Lightweight Transformer for On-Device Speech Emotion Recognition
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
von: Zhang, Zixing, et al.
Veröffentlicht: (2024)
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
von: Han, HyoJung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025) -
Learning Video Temporal Dynamics with Cross-Modal Attention for Robust Audio-Visual Speech Recognition
von: Kim, Sungnyun, et al.
Veröffentlicht: (2024) -
One-Class Learning with Adaptive Centroid Shift for Audio Deepfake Detection
von: Kim, Hyun Myung, et al.
Veröffentlicht: (2024) -
Two Heads Are Better Than One: Audio-Visual Speech Error Correction with Dual Hypotheses
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025) -
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
von: Wang, Yujin, et al.
Veröffentlicht: (2022)