Sustainable self-supervised learning for speech representations
Fuente:
arXiv
Salvato in:
| Autori principali: | Lugo, Luis, Vielzeuf, Valentin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Emergent morpho-phonological representations in self-supervised speech models
di: Gauthier, Jon, et al.
Pubblicazione: (2025)
di: Gauthier, Jon, et al.
Pubblicazione: (2025)
Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations
di: Mohamed, Mukhtar, et al.
Pubblicazione: (2024)
di: Mohamed, Mukhtar, et al.
Pubblicazione: (2024)
Investigating the 'Autoencoder Behavior' in Speech Self-Supervised Models: a focus on HuBERT's Pretraining
di: Vielzeuf, Valentin
Pubblicazione: (2024)
di: Vielzeuf, Valentin
Pubblicazione: (2024)
A low latency attention module for streaming self-supervised speech representation learning
di: Ma, Jianbo, et al.
Pubblicazione: (2023)
di: Ma, Jianbo, et al.
Pubblicazione: (2023)
Self-supervised learning of speech representations with Dutch archival data
di: Vaessen, Nik, et al.
Pubblicazione: (2025)
di: Vaessen, Nik, et al.
Pubblicazione: (2025)
AfriHuBERT: A self-supervised speech representation model for African languages
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
di: Fujita, Kenichi, et al.
Pubblicazione: (2024)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
Employing self-supervised learning models for cross-linguistic child speech maturity classification
di: Zhang, Theo, et al.
Pubblicazione: (2025)
di: Zhang, Theo, et al.
Pubblicazione: (2025)
Do self-supervised speech and language models extract similar representations as human brain?
di: Chen, Peili, et al.
Pubblicazione: (2023)
di: Chen, Peili, et al.
Pubblicazione: (2023)
Tracking the emergence of linguistic structure in self-supervised models learning from speech
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2026)
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2026)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
Probing self-attention in self-supervised speech models for cross-linguistic differences
di: Gopinath, Sai, et al.
Pubblicazione: (2024)
di: Gopinath, Sai, et al.
Pubblicazione: (2024)
Efficient infusion of self-supervised representations in Automatic Speech Recognition
di: Prabhu, Darshan, et al.
Pubblicazione: (2024)
di: Prabhu, Darshan, et al.
Pubblicazione: (2024)
Word stress in self-supervised speech models: A cross-linguistic comparison
di: Bentum, Martijn, et al.
Pubblicazione: (2025)
di: Bentum, Martijn, et al.
Pubblicazione: (2025)
Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
di: Herron, Felix, et al.
Pubblicazione: (2026)
di: Herron, Felix, et al.
Pubblicazione: (2026)
Investigating Low-Cost LLM Annotation for~Spoken Dialogue Understanding Datasets
di: Druart, Lucas, et al.
Pubblicazione: (2024)
di: Druart, Lucas, et al.
Pubblicazione: (2024)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
di: Gubian, Michele, et al.
Pubblicazione: (2025)
di: Gubian, Michele, et al.
Pubblicazione: (2025)
SENS-ASR: Semantic Embedding injection in Neural-transducer for Streaming Automatic Speech Recognition
di: Dkhissi, Youness, et al.
Pubblicazione: (2026)
di: Dkhissi, Youness, et al.
Pubblicazione: (2026)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
Is one brick enough to break the wall of spoken dialogue state tracking?
di: Druart, Lucas, et al.
Pubblicazione: (2023)
di: Druart, Lucas, et al.
Pubblicazione: (2023)
The Speech-LLM Takes It All: A Truly Fully End-to-End Spoken Dialogue State Tracking Approach
di: Ghazal, Nizar El, et al.
Pubblicazione: (2025)
di: Ghazal, Nizar El, et al.
Pubblicazione: (2025)
Semantic enrichment towards efficient speech representations
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2023)
di: Laperrière, Gaëlle, et al.
Pubblicazione: (2023)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2025)
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2025)
Self-supervised vision-langage alignment of deep learning representations for bone X-rays analysis
di: Englebert, Alexandre, et al.
Pubblicazione: (2024)
di: Englebert, Alexandre, et al.
Pubblicazione: (2024)
Hate speech detection in algerian dialect using deep learning
di: Lanasri, Dihia, et al.
Pubblicazione: (2023)
di: Lanasri, Dihia, et al.
Pubblicazione: (2023)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
di: Liu, Oli Danyi, et al.
Pubblicazione: (2024)
di: Liu, Oli Danyi, et al.
Pubblicazione: (2024)
The order in speech disorder: a scoping review of state of the art machine learning methods for clinical speech classification
di: Moell, Birger, et al.
Pubblicazione: (2025)
di: Moell, Birger, et al.
Pubblicazione: (2025)
Investigating the impact of 2D gesture representation on co-speech gesture generation
di: Guichoux, Teo, et al.
Pubblicazione: (2024)
di: Guichoux, Teo, et al.
Pubblicazione: (2024)
Learning representations of learning representations
di: González-Márquez, Rita, et al.
Pubblicazione: (2024)
di: González-Márquez, Rita, et al.
Pubblicazione: (2024)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
di: Garg, Abhinav, et al.
Pubblicazione: (2024)
di: Garg, Abhinav, et al.
Pubblicazione: (2024)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024)
di: Paraskevopoulos, Georgios, et al.
Pubblicazione: (2024)
Revisiting speech segmentation and lexicon learning with better features
di: Kamper, Herman, et al.
Pubblicazione: (2024)
di: Kamper, Herman, et al.
Pubblicazione: (2024)
Can Whisper perform speech-based in-context learning?
di: Wang, Siyin, et al.
Pubblicazione: (2023)
di: Wang, Siyin, et al.
Pubblicazione: (2023)
What does it take to get state of the art in simultaneous speech-to-speech translation?
di: Wilmet, Vincent, et al.
Pubblicazione: (2024)
di: Wilmet, Vincent, et al.
Pubblicazione: (2024)
Encoding of lexical tone in self-supervised models of spoken language
di: Shen, Gaofei, et al.
Pubblicazione: (2024)
di: Shen, Gaofei, et al.
Pubblicazione: (2024)
Cropping outperforms dropout as an augmentation strategy for self-supervised training of text embeddings
di: González-Márquez, Rita, et al.
Pubblicazione: (2025)
di: González-Márquez, Rita, et al.
Pubblicazione: (2025)
Linguists should learn to love speech-based deep learning models
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2025)
di: Kloots, Marianne de Heer, et al.
Pubblicazione: (2025)
On the limited utility of parallel data for learning shared multilingual representations
di: Leino, Julius, et al.
Pubblicazione: (2026)
di: Leino, Julius, et al.
Pubblicazione: (2026)
A self-supervised framework for learning whole slide representations
di: Hou, Xinhai, et al.
Pubblicazione: (2024)
di: Hou, Xinhai, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Emergent morpho-phonological representations in self-supervised speech models
di: Gauthier, Jon, et al.
Pubblicazione: (2025) -
Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations
di: Mohamed, Mukhtar, et al.
Pubblicazione: (2024) -
Investigating the 'Autoencoder Behavior' in Speech Self-Supervised Models: a focus on HuBERT's Pretraining
di: Vielzeuf, Valentin
Pubblicazione: (2024) -
A low latency attention module for streaming self-supervised speech representation learning
di: Ma, Jianbo, et al.
Pubblicazione: (2023) -
Self-supervised learning of speech representations with Dutch archival data
di: Vaessen, Nik, et al.
Pubblicazione: (2025)