Self-supervised learning of speech representations with Dutch archival data
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Vaessen, Nik, Ordelman, Roeland, van Leeuwen, David A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
von: Vaessen, Nik, et al.
Veröffentlicht: (2024)
A low latency attention module for streaming self-supervised speech representation learning
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
von: Ma, Jianbo, et al.
Veröffentlicht: (2023)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024)
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
AfriHuBERT: A self-supervised speech representation model for African languages
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
von: Alabi, Jesujoba O., et al.
Veröffentlicht: (2024)
Introduction to speech recognition
von: Dauphin, Gabriel
Veröffentlicht: (2024)
von: Dauphin, Gabriel
Veröffentlicht: (2024)
Acoustic-to-articulatory inversion for dysarthric speech: Are pre-trained self-supervised representations favorable?
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2023)
von: Maharana, Sarthak Kumar, et al.
Veröffentlicht: (2023)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
von: Gowda, Harshavardhana T., et al.
Veröffentlicht: (2025)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
von: Slabbert, Danel, et al.
Veröffentlicht: (2025)
Non-verbal information in spontaneous speech -- towards a new framework of analysis
von: Biron, Tirza, et al.
Veröffentlicht: (2024)
von: Biron, Tirza, et al.
Veröffentlicht: (2024)
Revisiting speech segmentation and lexicon learning with better features
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
von: Kamper, Herman, et al.
Veröffentlicht: (2024)
Semantic enrichment towards efficient speech representations
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
von: Laperrière, Gaëlle, et al.
Veröffentlicht: (2023)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
von: Paraskevopoulos, Georgios, et al.
Veröffentlicht: (2024)
A multimodal dynamical variational autoencoder for audiovisual speech representation learning
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
von: Sadok, Samir, et al.
Veröffentlicht: (2023)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
von: Liu, Oli Danyi, et al.
Veröffentlicht: (2024)
von: Liu, Oli Danyi, et al.
Veröffentlicht: (2024)
Throat and acoustic paired speech dataset for deep learning-based speech enhancement
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
von: Kim, Yunsik, et al.
Veröffentlicht: (2025)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
von: Kloots, Marianne de Heer, et al.
Veröffentlicht: (2025)
Selfsupervised learning for pathological speech detection
von: Sheikh, Shakeel Ahmad
Veröffentlicht: (2024)
von: Sheikh, Shakeel Ahmad
Veröffentlicht: (2024)
Pre-Trained Foundation Model representations to uncover Breathing patterns in Speech
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2024)
von: Mitra, Vikramjit, et al.
Veröffentlicht: (2024)
Data-driven grapheme-to-phoneme representations for a lexicon-free text-to-speech
von: Garg, Abhinav, et al.
Veröffentlicht: (2024)
von: Garg, Abhinav, et al.
Veröffentlicht: (2024)
Late fusion ensembles for speech recognition on diverse input audio representations
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
von: Jezidžić, Marin, et al.
Veröffentlicht: (2024)
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)
Can Whisper perform speech-based in-context learning?
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
von: Wang, Siyin, et al.
Veröffentlicht: (2023)
Task Oriented Dialogue as a Catalyst for Self-Supervised Automatic Speech Recognition
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
Improving Self-supervised Pre-training using Accent-Specific Codebooks
von: Prabhu, Darshan, et al.
Veröffentlicht: (2024)
von: Prabhu, Darshan, et al.
Veröffentlicht: (2024)
Self-consistent context aware conformer transducer for speech recognition
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
von: Kolokolov, Konstantin, et al.
Veröffentlicht: (2024)
Self-Train Before You Transcribe
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
von: Flynn, Robert, et al.
Veröffentlicht: (2024)
MERT: Acoustic Music Understanding Model with Large-Scale Self-supervised Training
von: Li, Yizhi, et al.
Veröffentlicht: (2023)
von: Li, Yizhi, et al.
Veröffentlicht: (2023)
Property Neurons in Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2024)
Towards Early Prediction of Self-Supervised Speech Model Performance
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
von: Whetten, Ryan, et al.
Veröffentlicht: (2025)
Is Smaller Always Faster? Tradeoffs in Compressing Self-Supervised Speech Transformers
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
von: Lin, Tzu-Quan, et al.
Veröffentlicht: (2022)
Guiding Frame-Level CTC Alignments Using Self-knowledge Distillation
von: Kim, Eungbeom, et al.
Veröffentlicht: (2024)
von: Kim, Eungbeom, et al.
Veröffentlicht: (2024)
Generalizable speech deepfake detection via meta-learned LoRA
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
von: Laakkonen, Janne, et al.
Veröffentlicht: (2025)
Efficient Training of Self-Supervised Speech Foundation Models on a Compute Budget
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
von: Liu, Andy T., et al.
Veröffentlicht: (2024)
SKILL: Similarity-aware Knowledge distILLation for Speech Self-Supervised Learning
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
von: Zampierin, Luca, et al.
Veröffentlicht: (2024)
Improving the Inclusivity of Dutch Speech Recognition by Fine-tuning Whisper on the JASMIN-CGN Corpus
von: Shekoufandeh, Golshid, et al.
Veröffentlicht: (2025)
von: Shekoufandeh, Golshid, et al.
Veröffentlicht: (2025)
Improving child speech recognition with augmented child-like speech
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuanyuan, et al.
Veröffentlicht: (2024)
CA-SSLR: Condition-Aware Self-Supervised Learning Representation for Generalized Speech Processing
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
von: Lu, Yen-Ju, et al.
Veröffentlicht: (2024)
Self-supervised Speech Representations Still Struggle with African American Vernacular English
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
von: Chang, Kalvin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
The Effect of Batch Size on Contrastive Self-Supervised Speech Representation Learning
von: Vaessen, Nik, et al.
Veröffentlicht: (2024) -
A low latency attention module for streaming self-supervised speech representation learning
von: Ma, Jianbo, et al.
Veröffentlicht: (2023) -
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
von: Wang, Hsuan-Fu, et al.
Veröffentlicht: (2024) -
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
von: Fujita, Kenichi, et al.
Veröffentlicht: (2024) -
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
von: Choi, Kwanghee, et al.
Veröffentlicht: (2026)