Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Alexander H., Yeh, Sung-Lin, Glass, James |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
USAD: Universal Speech and Audio Representation via Distillation
di: Chang, Heng-Jui, et al.
Pubblicazione: (2025)
di: Chang, Heng-Jui, et al.
Pubblicazione: (2025)
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
di: Liu, Alexander H., et al.
Pubblicazione: (2025)
di: Liu, Alexander H., et al.
Pubblicazione: (2025)
LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
di: Xue, Hongfei, et al.
Pubblicazione: (2023)
R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
di: Chang, Heng-Jui, et al.
Pubblicazione: (2023)
di: Chang, Heng-Jui, et al.
Pubblicazione: (2023)
Position-invariant Fine-tuning of Speech Enhancement Models with Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2026)
di: Meghanani, Amit, et al.
Pubblicazione: (2026)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
di: Chang, Kalvin, et al.
Pubblicazione: (2024)
di: Chang, Kalvin, et al.
Pubblicazione: (2024)
Improving Acoustic Word Embeddings through Correspondence Training of Self-supervised Speech Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
Benchmarking Children's ASR with Supervised and Self-supervised Speech Foundation Models
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
di: Fan, Ruchao, et al.
Pubblicazione: (2024)
Implicit Self-supervised Language Representation for Spoken Language Diarization
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
di: Mishra, Jagabandhu, et al.
Pubblicazione: (2023)
SCORE: Self-supervised Correspondence Fine-tuning for Improved Content Representations
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
di: Meghanani, Amit, et al.
Pubblicazione: (2024)
Rethinking Speech Representation Aggregation in Speech Enhancement: A Phonetic Mutual Information Perspective
di: Han, Seungu, et al.
Pubblicazione: (2026)
di: Han, Seungu, et al.
Pubblicazione: (2026)
Enhancing Few-shot Keyword Spotting Performance through Pre-Trained Self-supervised Speech Models
di: Gok, Alican, et al.
Pubblicazione: (2025)
di: Gok, Alican, et al.
Pubblicazione: (2025)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
di: Lin, Tzu-Quan, et al.
Pubblicazione: (2025)
Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?
di: Osakuade, Opeyemi, et al.
Pubblicazione: (2024)
di: Osakuade, Opeyemi, et al.
Pubblicazione: (2024)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
AV2Wav: Diffusion-Based Re-synthesis from Continuous Self-supervised Features for Audio-Visual Speech Enhancement
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2023)
di: Chou, Ju-Chieh, et al.
Pubblicazione: (2023)
Fast Word Error Rate Estimation Using Self-Supervised Representations for Speech and Text
di: Park, Chanho, et al.
Pubblicazione: (2023)
di: Park, Chanho, et al.
Pubblicazione: (2023)
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
di: Venkateswaran, Nitin, et al.
Pubblicazione: (2025)
di: Venkateswaran, Nitin, et al.
Pubblicazione: (2025)
Adapting Self-Supervised Speech Representations for Cross-lingual Dysarthria Detection in Parkinson's Disease
di: Hernandez, Abner, et al.
Pubblicazione: (2026)
di: Hernandez, Abner, et al.
Pubblicazione: (2026)
XLAVS-R: Cross-Lingual Audio-Visual Speech Representation Learning for Noise-Robust Speech Perception
di: Han, HyoJung, et al.
Pubblicazione: (2024)
di: Han, HyoJung, et al.
Pubblicazione: (2024)
Self-supervised Speech Models for Word-Level Stuttered Speech Detection
di: Shih, Yi-Jen, et al.
Pubblicazione: (2024)
di: Shih, Yi-Jen, et al.
Pubblicazione: (2024)
LeBenchmark 2.0: a Standardized, Replicable and Enhanced Framework for Self-supervised Representations of French Speech
di: Parcollet, Titouan, et al.
Pubblicazione: (2023)
di: Parcollet, Titouan, et al.
Pubblicazione: (2023)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
di: Peng, Junyi, et al.
Pubblicazione: (2025)
di: Peng, Junyi, et al.
Pubblicazione: (2025)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
di: Jang, Kangwook, et al.
Pubblicazione: (2023)
di: Jang, Kangwook, et al.
Pubblicazione: (2023)
A Closer Look at Neural Codec Resynthesis: Bridging the Gap between Codec and Waveform Generation
di: Liu, Alexander H., et al.
Pubblicazione: (2024)
di: Liu, Alexander H., et al.
Pubblicazione: (2024)
Conversational Speech Recognition by Learning Audio-textual Cross-modal Contextual Representation
di: Wei, Kun, et al.
Pubblicazione: (2023)
di: Wei, Kun, et al.
Pubblicazione: (2023)
Sylber: Syllabic Embedding Representation of Speech from Raw Audio
di: Cho, Cheol Jun, et al.
Pubblicazione: (2024)
di: Cho, Cheol Jun, et al.
Pubblicazione: (2024)
Speech Recognition Transformers: Topological-lingualism Perspective
di: Singh, Shruti, et al.
Pubblicazione: (2024)
di: Singh, Shruti, et al.
Pubblicazione: (2024)
Convexity-based Pruning of Speech Representation Models
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
di: Dorszewski, Teresa, et al.
Pubblicazione: (2024)
Configurable Multilingual ASR with Speech Summary Representations
di: Zhu, Harrison, et al.
Pubblicazione: (2024)
di: Zhu, Harrison, et al.
Pubblicazione: (2024)
Representation Purification for End-to-End Speech Translation
di: Zhang, Chengwei, et al.
Pubblicazione: (2024)
di: Zhang, Chengwei, et al.
Pubblicazione: (2024)
Enhancing Dialogue Speech Recognition with Robust Contextual Awareness via Noise Representation Learning
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
di: Lee, Wonjun, et al.
Pubblicazione: (2024)
Weight Factorization and Centralization for Continual Learning in Speech Recognition
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
di: Ugan, Enes Yavuz, et al.
Pubblicazione: (2025)
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
di: Li, Yuxin, et al.
Pubblicazione: (2025)
di: Li, Yuxin, et al.
Pubblicazione: (2025)
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2026)
di: Lin, Yi-Cheng, et al.
Pubblicazione: (2026)
Is Self-Supervised Learning Enough to Fill in the Gap? A Study on Speech Inpainting
di: Asaad, Ihab, et al.
Pubblicazione: (2024)
di: Asaad, Ihab, et al.
Pubblicazione: (2024)
Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
di: Phukan, Orchid Chetia, et al.
Pubblicazione: (2024)
Rethinking Discrete Speech Representation Tokens for Accent Generation
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2026)
di: Zhong, Jinzuomu, et al.
Pubblicazione: (2026)
Task-Agnostic Structured Pruning of Speech Representation Models
di: Wang, Haoyu, et al.
Pubblicazione: (2023)
di: Wang, Haoyu, et al.
Pubblicazione: (2023)
Documenti analoghi
-
USAD: Universal Speech and Audio Representation via Distillation
di: Chang, Heng-Jui, et al.
Pubblicazione: (2025) -
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
di: Liu, Alexander H., et al.
Pubblicazione: (2025) -
LASER: Learning by Aligning Self-supervised Representations of Speech for Improving Content-related Tasks
di: Meghanani, Amit, et al.
Pubblicazione: (2024) -
SSHR: Leveraging Self-supervised Hierarchical Representations for Multilingual Automatic Speech Recognition
di: Xue, Hongfei, et al.
Pubblicazione: (2023) -
R-Spin: Efficient Speaker and Noise-invariant Representation Learning with Acoustic Pieces
di: Chang, Heng-Jui, et al.
Pubblicazione: (2023)