Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Huo, Robin, Dunbar, Ewan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
von: Yoon, Ji Won, et al.
Veröffentlicht: (2022)
von: Yoon, Ji Won, et al.
Veröffentlicht: (2022)
Automatic classification of stop realisation with wav2vec2.0
von: Tanner, James, et al.
Veröffentlicht: (2025)
von: Tanner, James, et al.
Veröffentlicht: (2025)
mHuBERT-147: A Compact Multilingual HuBERT Model
von: Boito, Marcely Zanon, et al.
Veröffentlicht: (2024)
von: Boito, Marcely Zanon, et al.
Veröffentlicht: (2024)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
von: Bayerl, Sebastian P., et al.
Veröffentlicht: (2022)
von: Bayerl, Sebastian P., et al.
Veröffentlicht: (2022)
MelHuBERT: A simplified HuBERT on Mel spectrograms
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024)
von: Komatsu, Ryota, et al.
Veröffentlicht: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
von: Li, Xingyuan, et al.
Veröffentlicht: (2024)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
von: Wu, Wenxuan, et al.
Veröffentlicht: (2024)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
von: Wang, Zhiyong, et al.
Veröffentlicht: (2024)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
von: Cho, Cheol Jun, et al.
Veröffentlicht: (2023)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
von: Guo, Yiwei, et al.
Veröffentlicht: (2024)
RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
von: Chen, Xueyuan, et al.
Veröffentlicht: (2024)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
von: Shi, Jiatong, et al.
Veröffentlicht: (2023)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
von: Chi, Hyung Gun, et al.
Veröffentlicht: (2025)
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
von: Zhu, Qiushi, et al.
Veröffentlicht: (2024)
von: Zhu, Qiushi, et al.
Veröffentlicht: (2024)
Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0
von: Wang, Yueguan, et al.
Veröffentlicht: (2025)
von: Wang, Yueguan, et al.
Veröffentlicht: (2025)
AfriHuBERT: A self-supervised speech representation model for African languages
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
VisG AV-HuBERT: Viseme-Guided AV-HuBERT
von: Papadopoulos, Aristeidis, et al.
Veröffentlicht: (2026)
von: Papadopoulos, Aristeidis, et al.
Veröffentlicht: (2026)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
von: Li, Feng, et al.
Veröffentlicht: (2024)
von: Li, Feng, et al.
Veröffentlicht: (2024)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
von: Shao, Qijie, et al.
Veröffentlicht: (2025)
The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
von: Ong, Michael, et al.
Veröffentlicht: (2024)
von: Ong, Michael, et al.
Veröffentlicht: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
von: Attia, Ahmed Adel, et al.
Veröffentlicht: (2024)
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
von: Abdullah, Abdulhady Abas, et al.
Veröffentlicht: (2025)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
von: Robertson, Sean, et al.
Veröffentlicht: (2023)
Experimental Study: Enhancing Voice Spoofing Detection Models with wav2vec 2.0
von: Kang, Taein, et al.
Veröffentlicht: (2024)
von: Kang, Taein, et al.
Veröffentlicht: (2024)
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
von: Poli, Maxime, et al.
Veröffentlicht: (2026)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
von: Huang, Kuan-Po, et al.
Veröffentlicht: (2023)
BERT-LID: Leveraging BERT to Improve Spoken Language Identification
von: Nie, Yuting, et al.
Veröffentlicht: (2022)
von: Nie, Yuting, et al.
Veröffentlicht: (2022)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
von: Ahn, Hyebin, et al.
Veröffentlicht: (2025)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
von: Jafarzadeh, Pourya, et al.
Veröffentlicht: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025)
von: Rackauckas, Zackary, et al.
Veröffentlicht: (2025)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
von: Anidjar, Or Haim, et al.
Veröffentlicht: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
von: Grossman, Raymond, et al.
Veröffentlicht: (2025)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
von: Shi, Jiatong, et al.
Veröffentlicht: (2024)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
von: Cheng, Shanbo, et al.
Veröffentlicht: (2025)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
von: Wang, Qingzheng, et al.
Veröffentlicht: (2025)
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
von: Li, Ze, et al.
Veröffentlicht: (2025)
von: Li, Ze, et al.
Veröffentlicht: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
von: Yamauchi, Kazuki, et al.
Veröffentlicht: (2024)
Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration
von: Karita, Shigeki, et al.
Veröffentlicht: (2025)
von: Karita, Shigeki, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
von: Yoon, Ji Won, et al.
Veröffentlicht: (2022) -
Automatic classification of stop realisation with wav2vec2.0
von: Tanner, James, et al.
Veröffentlicht: (2025) -
mHuBERT-147: A Compact Multilingual HuBERT Model
von: Boito, Marcely Zanon, et al.
Veröffentlicht: (2024) -
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
von: Bayerl, Sebastian P., et al.
Veröffentlicht: (2022) -
MelHuBERT: A simplified HuBERT on Mel spectrograms
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)