Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
Fuente:
arXiv
Saved in:
| Main Authors: | Gubian, Michele, Krehan, Ioana, Liu, Oli, Kirby, James, Goldwater, Sharon |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
by: Liu, Oli Danyi, et al.
Published: (2024)
by: Liu, Oli Danyi, et al.
Published: (2024)
Curriculum learning for self-supervised speaker verification
by: Heo, Hee-Soo, et al.
Published: (2022)
by: Heo, Hee-Soo, et al.
Published: (2022)
Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations
by: Mohamed, Mukhtar, et al.
Published: (2024)
by: Mohamed, Mukhtar, et al.
Published: (2024)
Analyzing phonetic structure of Mandarin using Audacity
by: Xu, Shizheng
Published: (2024)
by: Xu, Shizheng
Published: (2024)
On the social bias of speech self-supervised models
by: Lin, Yi-Cheng, et al.
Published: (2024)
by: Lin, Yi-Cheng, et al.
Published: (2024)
End-to-end transfer learning for speaker-independent cross-language and cross-corpus speech emotion recognition
by: Tang, Duowei, et al.
Published: (2023)
by: Tang, Duowei, et al.
Published: (2023)
CardioPHON: Quality assessment and self-supervised pretraining for screening of cardiac function based on phonocardiogram recordings
by: Despotovic, Vladimir, et al.
Published: (2025)
by: Despotovic, Vladimir, et al.
Published: (2025)
Lightweight speech enhancement guided target speech extraction in noisy multi-speaker scenarios
by: Huang, Ziling, et al.
Published: (2025)
by: Huang, Ziling, et al.
Published: (2025)
AfriHuBERT: A self-supervised speech representation model for African languages
by: Alabi, Jesujoba O., et al.
Published: (2024)
by: Alabi, Jesujoba O., et al.
Published: (2024)
Evaluating pretrained speech embedding systems for dysarthria detection across heterogenous datasets
by: Wihlborg, Lovisa, et al.
Published: (2025)
by: Wihlborg, Lovisa, et al.
Published: (2025)
On the relationship between speech and hearing
by: Umesh, Srinivasan, et al.
Published: (2024)
by: Umesh, Srinivasan, et al.
Published: (2024)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
by: Kloots, Marianne de Heer, et al.
Published: (2025)
by: Kloots, Marianne de Heer, et al.
Published: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
by: Gowda, Harshavardhana T., et al.
Published: (2025)
by: Gowda, Harshavardhana T., et al.
Published: (2025)
SLM-S2ST: A multimodal language model for direct speech-to-speech translation
by: Hu, Yuxuan, et al.
Published: (2025)
by: Hu, Yuxuan, et al.
Published: (2025)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
by: Pascu, Octavian, et al.
Published: (2023)
by: Pascu, Octavian, et al.
Published: (2023)
Efficient training strategies for natural sounding speech synthesis and speaker adaptation based on FastPitch
by: Răgman, Teodora, et al.
Published: (2024)
by: Răgman, Teodora, et al.
Published: (2024)
Encoding of lexical tone in self-supervised models of spoken language
by: Shen, Gaofei, et al.
Published: (2024)
by: Shen, Gaofei, et al.
Published: (2024)
Effective Context in Neural Speech Models
by: Meng, Yen, et al.
Published: (2025)
by: Meng, Yen, et al.
Published: (2025)
Human-CLAP: Human-perception-based contrastive language-audio pretraining
by: Takano, Taisei, et al.
Published: (2025)
by: Takano, Taisei, et al.
Published: (2025)
Do self-supervised speech and language models extract similar representations as human brain?
by: Chen, Peili, et al.
Published: (2023)
by: Chen, Peili, et al.
Published: (2023)
On the influence of language similarity in non-target speaker verification trials
by: Reuter, Paul M., et al.
Published: (2025)
by: Reuter, Paul M., et al.
Published: (2025)
Spoken language change detection inspired by speaker change detection
by: Mishra, Jagabandhu, et al.
Published: (2023)
by: Mishra, Jagabandhu, et al.
Published: (2023)
GDiffuSE: Diffusion-based speech enhancement with noise model guidance
by: Yanir, Efrayim, et al.
Published: (2025)
by: Yanir, Efrayim, et al.
Published: (2025)
Thinking in cocktail party: Chain-of-Thought and reinforcement learning for target speaker automatic speech recognition
by: Zhang, Yiru, et al.
Published: (2025)
by: Zhang, Yiru, et al.
Published: (2025)
Spectral or spatial? Leveraging both for speaker extraction in challenging data conditions
by: Eisenberg, Aviad, et al.
Published: (2025)
by: Eisenberg, Aviad, et al.
Published: (2025)
Challenging margin-based speaker embedding extractors by using the variational information bottleneck
by: Stafylakis, Themos, et al.
Published: (2024)
by: Stafylakis, Themos, et al.
Published: (2024)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
by: Paraskevopoulos, Georgios, et al.
Published: (2024)
by: Paraskevopoulos, Georgios, et al.
Published: (2024)
The importance of spatial and spectral information in multiple speaker tracking
by: Beit-On, Hanan, et al.
Published: (2024)
by: Beit-On, Hanan, et al.
Published: (2024)
Probing mental health information in speech foundation models
by: de Gennes, Marc, et al.
Published: (2024)
by: de Gennes, Marc, et al.
Published: (2024)
Online incremental learning for audio classification using a pretrained audio model
by: Mulimani, Manjunath, et al.
Published: (2025)
by: Mulimani, Manjunath, et al.
Published: (2025)
Online speaker diarization of meetings guided by speech separation
by: Gruttadauria, Elio, et al.
Published: (2024)
by: Gruttadauria, Elio, et al.
Published: (2024)
Tracking the emergence of linguistic structure in self-supervised models learning from speech
by: Kloots, Marianne de Heer, et al.
Published: (2026)
by: Kloots, Marianne de Heer, et al.
Published: (2026)
Multi-speaker Text-to-speech Training with Speaker Anonymized Data
by: Huang, Wen-Chin, et al.
Published: (2024)
by: Huang, Wen-Chin, et al.
Published: (2024)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models
by: Jung, Jee-weon, et al.
Published: (2024)
by: Jung, Jee-weon, et al.
Published: (2024)
Transcribe, Align and Segment: Creating speech datasets for low-resource languages
by: Sereda, Taras
Published: (2024)
by: Sereda, Taras
Published: (2024)
Phoneme-based speech recognition driven by large language models and sampling marginalization
by: Ma, Te, et al.
Published: (2025)
by: Ma, Te, et al.
Published: (2025)
Hierarchical speaker representation for target speaker extraction
by: He, Shulin, et al.
Published: (2022)
by: He, Shulin, et al.
Published: (2022)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
by: Ogun, Sewade, et al.
Published: (2024)
by: Ogun, Sewade, et al.
Published: (2024)
Fine-tune the pretrained ATST model for sound event detection
by: Shao, Nian, et al.
Published: (2023)
by: Shao, Nian, et al.
Published: (2023)
Similar Items
-
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
by: Wang, Yi, et al.
Published: (2025) -
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
by: Liu, Oli Danyi, et al.
Published: (2024) -
Curriculum learning for self-supervised speaker verification
by: Heo, Hee-Soo, et al.
Published: (2022) -
Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations
by: Mohamed, Mukhtar, et al.
Published: (2024) -
Analyzing phonetic structure of Mandarin using Audacity
by: Xu, Shizheng
Published: (2024)