DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chi, Hyung Gun, Aldeneh, Zakaria, Likhomanenko, Tatiana, Rudovic, Oggi, Higuchi, Takuya, Chen, Li-Wei, Watanabe, Shinji, Abdelaziz, Ahmed Hussen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VisG AV-HuBERT: Viseme-Guided AV-HuBERT
von: Papadopoulos, Aristeidis, et al.
Veröffentlicht: (2026)
von: Papadopoulos, Aristeidis, et al.
Veröffentlicht: (2026)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
von: Yoon, Ji Won, et al.
Veröffentlicht: (2022)
von: Yoon, Ji Won, et al.
Veröffentlicht: (2022)
mHuBERT-147: A Compact Multilingual HuBERT Model
von: Boito, Marcely Zanon, et al.
Veröffentlicht: (2024)
von: Boito, Marcely Zanon, et al.
Veröffentlicht: (2024)
MelHuBERT: A simplified HuBERT on Mel spectrograms
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Can you Remove the Downstream Model for Speaker Recognition with Self-Supervised Speech Features?
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024)
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024)
Text-guided HuBERT: Self-Supervised Speech Pre-training via Generative Adversarial Networks
von: Ma, Duo, et al.
Veröffentlicht: (2024)
von: Ma, Duo, et al.
Veröffentlicht: (2024)
Towards Automatic Assessment of Self-Supervised Speech Models using Rank
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Adaptive Knowledge Distillation for Device-Directed Speech Detection
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
Exploring Prediction Targets in Masked Pre-Training for Speech Foundation Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2024)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
Speaker-IPL: Unsupervised Learning of Speaker Characteristics with i-Vector based Pseudo-Labels
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
von: Aldeneh, Zakaria, et al.
Veröffentlicht: (2024)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
von: Huo, Robin, et al.
Veröffentlicht: (2025)
von: Huo, Robin, et al.
Veröffentlicht: (2025)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
von: Jung, Jee-weon, et al.
Veröffentlicht: (2024)
A Variational Framework for Improving Naturalness in Generative Spoken Language Models
von: Chen, Li-Wei, et al.
Veröffentlicht: (2025)
von: Chen, Li-Wei, et al.
Veröffentlicht: (2025)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
BabyHuBERT: Multilingual Self-Supervised Learning for Segmenting Speakers in Child-Centered Long-Form Recordings
von: Charlot, Théo, et al.
Veröffentlicht: (2025)
von: Charlot, Théo, et al.
Veröffentlicht: (2025)
AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2023)
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2023)
AfriHuBERT: A self-supervised speech representation model for African languages
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
dMel: Speech Tokenization made Simple
von: Bai, Richard He, et al.
Veröffentlicht: (2024)
von: Bai, Richard He, et al.
Veröffentlicht: (2024)
Multimodal Large Language Models with Fusion Low Rank Adaptation for Device Directed Speech Detection
von: Palaskar, Shruti, et al.
Veröffentlicht: (2024)
von: Palaskar, Shruti, et al.
Veröffentlicht: (2024)
Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
von: Blandón, María Andrea Cruz, et al.
Veröffentlicht: (2025)
von: Blandón, María Andrea Cruz, et al.
Veröffentlicht: (2025)
Mechanisms of Multimodal Synchronization: Insights from Decoder-Based Video-Text-to-Speech Synthesis
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
von: Gupta, Akshita, et al.
Veröffentlicht: (2024)
Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
von: Narain, Jaya, et al.
Veröffentlicht: (2025)
Closing the Gap Between Text and Speech Understanding in LLMs
von: Cuervo, Santiago, et al.
Veröffentlicht: (2025)
von: Cuervo, Santiago, et al.
Veröffentlicht: (2025)
End-to-End Speech Recognition with Pre-trained Masked Language Model
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
von: Higuchi, Yosuke, et al.
Veröffentlicht: (2024)
HuPER: A Human-Inspired Framework for Phonetic Perception
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
von: Guo, Chenxu, et al.
Veröffentlicht: (2026)
BERT-LID: Leveraging BERT to Improve Spoken Language Identification
von: Nie, Yuting, et al.
Veröffentlicht: (2022)
von: Nie, Yuting, et al.
Veröffentlicht: (2022)
Segment-Level Vectorized Beam Search Based on Partially Autoregressive Inference
von: Someki, Masao, et al.
Veröffentlicht: (2023)
von: Someki, Masao, et al.
Veröffentlicht: (2023)
ChipChat: Low-Latency Cascaded Conversational Agent in MLX
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2025)
von: Likhomanenko, Tatiana, et al.
Veröffentlicht: (2025)
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
von: Li, Ze, et al.
Veröffentlicht: (2025)
von: Li, Ze, et al.
Veröffentlicht: (2025)
Prosody Labeling with Phoneme-BERT and Speech Foundation Models
von: Koriyama, Tomoki
Veröffentlicht: (2025)
von: Koriyama, Tomoki
Veröffentlicht: (2025)
Comparative Analysis of Personalized Voice Activity Detection Systems: Assessing Real-World Effectiveness
von: Kumar, Satyam, et al.
Veröffentlicht: (2024)
von: Kumar, Satyam, et al.
Veröffentlicht: (2024)
MMT-BERT: Chord-aware Symbolic Music Generation Based on Multitrack Music Transformer and MusicBERT
von: Zhu, Jinlong, et al.
Veröffentlicht: (2024)
von: Zhu, Jinlong, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VisG AV-HuBERT: Viseme-Guided AV-HuBERT
von: Papadopoulos, Aristeidis, et al.
Veröffentlicht: (2026) -
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023) -
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
von: Yoon, Ji Won, et al.
Veröffentlicht: (2022) -
mHuBERT-147: A Compact Multilingual HuBERT Model
von: Boito, Marcely Zanon, et al.
Veröffentlicht: (2024) -
MelHuBERT: A simplified HuBERT on Mel spectrograms
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)