Iterative refinement, not training objective, makes HuBERT behave differently from wav2vec 2.0
Fuente:
arXiv
Salvato in:
| Autori principali: | Huo, Robin, Dunbar, Ewan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
di: Yoon, Ji Won, et al.
Pubblicazione: (2022)
di: Yoon, Ji Won, et al.
Pubblicazione: (2022)
Automatic classification of stop realisation with wav2vec2.0
di: Tanner, James, et al.
Pubblicazione: (2025)
di: Tanner, James, et al.
Pubblicazione: (2025)
mHuBERT-147: A Compact Multilingual HuBERT Model
di: Boito, Marcely Zanon, et al.
Pubblicazione: (2024)
di: Boito, Marcely Zanon, et al.
Pubblicazione: (2024)
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
di: Bayerl, Sebastian P., et al.
Pubblicazione: (2022)
di: Bayerl, Sebastian P., et al.
Pubblicazione: (2022)
MelHuBERT: A simplified HuBERT on Mel spectrograms
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
di: Komatsu, Ryota, et al.
Pubblicazione: (2024)
di: Komatsu, Ryota, et al.
Pubblicazione: (2024)
Phonetic and Lexical Discovery of a Canine Language using HuBERT
di: Li, Xingyuan, et al.
Pubblicazione: (2024)
di: Li, Xingyuan, et al.
Pubblicazione: (2024)
Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover Strategy
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
di: Wu, Wenxuan, et al.
Pubblicazione: (2024)
Mixture of Experts Fusion for Fake Audio Detection Using Frozen wav2vec 2.0
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
di: Wang, Zhiyong, et al.
Pubblicazione: (2024)
SD-HuBERT: Sentence-Level Self-Distillation Induces Syllabic Organization in HuBERT
di: Cho, Cheol Jun, et al.
Pubblicazione: (2023)
di: Cho, Cheol Jun, et al.
Pubblicazione: (2023)
vec2wav 2.0: Advancing Voice Conversion via Discrete Token Vocoders
di: Guo, Yiwei, et al.
Pubblicazione: (2024)
di: Guo, Yiwei, et al.
Pubblicazione: (2024)
RobustSVC: HuBERT-based Melody Extractor and Adversarial Learning for Robust Singing Voice Conversion
di: Chen, Wei, et al.
Pubblicazione: (2024)
di: Chen, Wei, et al.
Pubblicazione: (2024)
Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction
di: Chen, Xueyuan, et al.
Pubblicazione: (2024)
di: Chen, Xueyuan, et al.
Pubblicazione: (2024)
Multi-resolution HuBERT: Multi-resolution Speech Self-Supervised Learning with Masked Unit Prediction
di: Shi, Jiatong, et al.
Pubblicazione: (2023)
di: Shi, Jiatong, et al.
Pubblicazione: (2023)
DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective
di: Chi, Hyung Gun, et al.
Pubblicazione: (2025)
di: Chi, Hyung Gun, et al.
Pubblicazione: (2025)
Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation
di: Zhu, Qiushi, et al.
Pubblicazione: (2024)
di: Zhu, Qiushi, et al.
Pubblicazione: (2024)
Enhancing and Exploring Mild Cognitive Impairment Detection with W2V-BERT-2.0
di: Wang, Yueguan, et al.
Pubblicazione: (2025)
di: Wang, Yueguan, et al.
Pubblicazione: (2025)
AfriHuBERT: A self-supervised speech representation model for African languages
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
VisG AV-HuBERT: Viseme-Guided AV-HuBERT
di: Papadopoulos, Aristeidis, et al.
Pubblicazione: (2026)
di: Papadopoulos, Aristeidis, et al.
Pubblicazione: (2026)
WavFusion: Towards wav2vec 2.0 Multimodal Speech Emotion Recognition
di: Li, Feng, et al.
Pubblicazione: (2024)
di: Li, Feng, et al.
Pubblicazione: (2024)
DQ-Data2vec: Decoupling Quantization for Multilingual Speech Recognition
di: Shao, Qijie, et al.
Pubblicazione: (2025)
di: Shao, Qijie, et al.
Pubblicazione: (2025)
The Faetar Benchmark: Speech Recognition in a Very Under-Resourced Language
di: Ong, Michael, et al.
Pubblicazione: (2024)
di: Ong, Michael, et al.
Pubblicazione: (2024)
CPT-Boosted Wav2vec2.0: Towards Noise Robust Speech Recognition for Classroom Environments
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
di: Attia, Ahmed Adel, et al.
Pubblicazione: (2024)
Speaker Diarization for Low-Resource Languages Through Wav2vec Fine-Tuning
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2025)
di: Abdullah, Abdulhady Abas, et al.
Pubblicazione: (2025)
Bigger is not Always Better: The Effect of Context Size on Speech Pre-Training
di: Robertson, Sean, et al.
Pubblicazione: (2023)
di: Robertson, Sean, et al.
Pubblicazione: (2023)
Experimental Study: Enhancing Voice Spoofing Detection Models with wav2vec 2.0
di: Kang, Taein, et al.
Pubblicazione: (2024)
di: Kang, Taein, et al.
Pubblicazione: (2024)
DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units
di: Poli, Maxime, et al.
Pubblicazione: (2026)
di: Poli, Maxime, et al.
Pubblicazione: (2026)
Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
di: Huang, Kuan-Po, et al.
Pubblicazione: (2023)
BERT-LID: Leveraging BERT to Improve Spoken Language Identification
di: Nie, Yuting, et al.
Pubblicazione: (2022)
di: Nie, Yuting, et al.
Pubblicazione: (2022)
HuBERT-VIC: Improving Noise-Robust Automatic Speech Recognition of Speech Foundation Model via Variance-Invariance-Covariance Regularization
di: Ahn, Hyebin, et al.
Pubblicazione: (2025)
di: Ahn, Hyebin, et al.
Pubblicazione: (2025)
Speaker Emotion Recognition: Leveraging Self-Supervised Models for Feature Extraction Using Wav2Vec2 and HuBERT
di: Jafarzadeh, Pourya, et al.
Pubblicazione: (2024)
di: Jafarzadeh, Pourya, et al.
Pubblicazione: (2024)
Comparative Evaluation of Expressive Japanese Character Text-to-Speech with VITS and Style-BERT-VITS2
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
di: Rackauckas, Zackary, et al.
Pubblicazione: (2025)
Whisper Turns Stronger: Augmenting Wav2Vec 2.0 for Superior ASR in Low-Resource Languages
di: Anidjar, Or Haim, et al.
Pubblicazione: (2024)
di: Anidjar, Or Haim, et al.
Pubblicazione: (2024)
SPGISpeech 2.0: Transcribed multi-speaker financial audio for speaker-tagged transcription
di: Grossman, Raymond, et al.
Pubblicazione: (2025)
di: Grossman, Raymond, et al.
Pubblicazione: (2025)
ML-SUPERB 2.0: Benchmarking Multilingual Speech Models Across Modeling Constraints, Languages, and Datasets
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
di: Shi, Jiatong, et al.
Pubblicazione: (2024)
Seed LiveInterpret 2.0: End-to-end Simultaneous Speech-to-speech Translation with Your Voice
di: Cheng, Shanbo, et al.
Pubblicazione: (2025)
di: Cheng, Shanbo, et al.
Pubblicazione: (2025)
Improving Multilingual Speech Models on ML-SUPERB 2.0: Fine-tuning with Data Augmentation and LID-Aware CTC
di: Wang, Qingzheng, et al.
Pubblicazione: (2025)
di: Wang, Qingzheng, et al.
Pubblicazione: (2025)
Enhancing Speaker Verification with w2v-BERT 2.0 and Knowledge Distillation guided Structured Pruning
di: Li, Ze, et al.
Pubblicazione: (2025)
di: Li, Ze, et al.
Pubblicazione: (2025)
Cross-Dialect Text-To-Speech in Pitch-Accent Language Incorporating Multi-Dialect Phoneme-Level BERT
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
di: Yamauchi, Kazuki, et al.
Pubblicazione: (2024)
Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration
di: Karita, Shigeki, et al.
Pubblicazione: (2025)
di: Karita, Shigeki, et al.
Pubblicazione: (2025)
Documenti analoghi
-
HuBERT-EE: Early Exiting HuBERT for Efficient Speech Recognition
di: Yoon, Ji Won, et al.
Pubblicazione: (2022) -
Automatic classification of stop realisation with wav2vec2.0
di: Tanner, James, et al.
Pubblicazione: (2025) -
mHuBERT-147: A Compact Multilingual HuBERT Model
di: Boito, Marcely Zanon, et al.
Pubblicazione: (2024) -
Detecting Dysfluencies in Stuttering Therapy Using wav2vec 2.0
di: Bayerl, Sebastian P., et al.
Pubblicazione: (2022) -
MelHuBERT: A simplified HuBERT on Mel spectrograms
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2022)