Orthogonality and isotropy of speaker and phonetic information in self-supervised speech representations
Fuente:
arXiv
Saved in:
| Main Authors: | Mohamed, Mukhtar, Liu, Oli Danyi, Tang, Hao, Goldwater, Sharon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
by: Gubian, Michele, et al.
Published: (2025)
by: Gubian, Michele, et al.
Published: (2025)
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
by: Liu, Oli Danyi, et al.
Published: (2024)
by: Liu, Oli Danyi, et al.
Published: (2024)
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
A framework for analyzing concept representations in neural models
by: Naowarat, Burin, et al.
Published: (2026)
by: Naowarat, Burin, et al.
Published: (2026)
Sustainable self-supervised learning for speech representations
by: Lugo, Luis, et al.
Published: (2024)
by: Lugo, Luis, et al.
Published: (2024)
Emergent morpho-phonological representations in self-supervised speech models
by: Gauthier, Jon, et al.
Published: (2025)
by: Gauthier, Jon, et al.
Published: (2025)
Effective Context in Neural Speech Models
by: Meng, Yen, et al.
Published: (2025)
by: Meng, Yen, et al.
Published: (2025)
Code-switching in text and speech challenges information-theoretic speaker design
by: Bhattacharya, Debasmita, et al.
Published: (2024)
by: Bhattacharya, Debasmita, et al.
Published: (2024)
AfriHuBERT: A self-supervised speech representation model for African languages
by: Alabi, Jesujoba O., et al.
Published: (2024)
by: Alabi, Jesujoba O., et al.
Published: (2024)
Revisiting Common Assumptions about Arabic Dialects in NLP
by: Keleg, Amr, et al.
Published: (2025)
by: Keleg, Amr, et al.
Published: (2025)
A Grounded Typology of Word Classes
by: Haley, Coleman, et al.
Published: (2024)
by: Haley, Coleman, et al.
Published: (2024)
Estimating the Level of Dialectness Predicts Interannotator Agreement in Multi-dialect Arabic Datasets
by: Keleg, Amr, et al.
Published: (2024)
by: Keleg, Amr, et al.
Published: (2024)
Noise-robust zero-shot text-to-speech synthesis conditioned on self-supervised speech-representation model with adapters
by: Fujita, Kenichi, et al.
Published: (2024)
by: Fujita, Kenichi, et al.
Published: (2024)
A low latency attention module for streaming self-supervised speech representation learning
by: Ma, Jianbo, et al.
Published: (2023)
by: Ma, Jianbo, et al.
Published: (2023)
Do self-supervised speech and language models extract similar representations as human brain?
by: Chen, Peili, et al.
Published: (2023)
by: Chen, Peili, et al.
Published: (2023)
A stylometric analysis of speaker attribution from speech transcripts
by: Aggazzotti, Cristina, et al.
Published: (2025)
by: Aggazzotti, Cristina, et al.
Published: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
by: Gowda, Harshavardhana T., et al.
Published: (2025)
by: Gowda, Harshavardhana T., et al.
Published: (2025)
Probing self-attention in self-supervised speech models for cross-linguistic differences
by: Gopinath, Sai, et al.
Published: (2024)
by: Gopinath, Sai, et al.
Published: (2024)
On the performance of phonetic algorithms in microtext normalization
by: Doval, Yerai, et al.
Published: (2024)
by: Doval, Yerai, et al.
Published: (2024)
Efficient infusion of self-supervised representations in Automatic Speech Recognition
by: Prabhu, Darshan, et al.
Published: (2024)
by: Prabhu, Darshan, et al.
Published: (2024)
Self-supervised learning of speech representations with Dutch archival data
by: Vaessen, Nik, et al.
Published: (2025)
by: Vaessen, Nik, et al.
Published: (2025)
Word stress in self-supervised speech models: A cross-linguistic comparison
by: Bentum, Martijn, et al.
Published: (2025)
by: Bentum, Martijn, et al.
Published: (2025)
Identifying and typifying demographic unfairness in phoneme-level embeddings of self-supervised speech recognition models
by: Herron, Felix, et al.
Published: (2026)
by: Herron, Felix, et al.
Published: (2026)
Addressing speaker gender bias in large scale speech translation systems
by: Bansal, Shubham, et al.
Published: (2025)
by: Bansal, Shubham, et al.
Published: (2025)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
by: Wang, Hsuan-Fu, et al.
Published: (2024)
by: Wang, Hsuan-Fu, et al.
Published: (2024)
Employing self-supervised learning models for cross-linguistic child speech maturity classification
by: Zhang, Theo, et al.
Published: (2025)
by: Zhang, Theo, et al.
Published: (2025)
Tracking the emergence of linguistic structure in self-supervised models learning from speech
by: Kloots, Marianne de Heer, et al.
Published: (2026)
by: Kloots, Marianne de Heer, et al.
Published: (2026)
The agreement of phonetic transcriptions between paediatric speech and language therapists transcribing a disordered speech sample
by: Laura Jane Mallaband
Published: (2024)
by: Laura Jane Mallaband
Published: (2024)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
by: Ogun, Sewade, et al.
Published: (2024)
by: Ogun, Sewade, et al.
Published: (2024)
Learner training for phonetic transcription of typical and/or disordered speech: A scoping review
by: Alice Lee, et al.
Published: (2024)
by: Alice Lee, et al.
Published: (2024)
Cross-linguistically Consistent Semantic and Syntactic Annotation of Child-directed Speech
by: Szubert, Ida, et al.
Published: (2021)
by: Szubert, Ida, et al.
Published: (2021)
Parallel Needleman-Wunsch on CUDA to measure word similarity based on phonetic transcriptions
by: Plein, Dominic
Published: (2025)
by: Plein, Dominic
Published: (2025)
Retrieval or Representation? Reassessing Benchmark Gaps in Multilingual and Visually Rich RAG
by: Asenov, Martin, et al.
Published: (2026)
by: Asenov, Martin, et al.
Published: (2026)
PyPhonPlan: Simulating phonetic planning with dynamic neural fields and task dynamics
by: Kirkham, Sam
Published: (2026)
by: Kirkham, Sam
Published: (2026)
Target speaker anonymization in multi-speaker recordings
by: Tomashenko, Natalia, et al.
Published: (2025)
by: Tomashenko, Natalia, et al.
Published: (2025)
The Development of a Comprehensive Spanish Dictionary for Phonetic and Lexical Tagging in Socio-phonetic Research (ESPADA)
by: Gonzalez, Simon
Published: (2024)
by: Gonzalez, Simon
Published: (2024)
Semantic enrichment towards efficient speech representations
by: Laperrière, Gaëlle, et al.
Published: (2023)
by: Laperrière, Gaëlle, et al.
Published: (2023)
Exploring the topics, sentiments and hate speech in the Spanish information environment
by: LOPEZ, ALEJANDRO BUITRAGO, et al.
Published: (2024)
by: LOPEZ, ALEJANDRO BUITRAGO, et al.
Published: (2024)
Extending Whisper with prompt tuning to target-speaker ASR
by: Ma, Hao, et al.
Published: (2023)
by: Ma, Hao, et al.
Published: (2023)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
by: Kloots, Marianne de Heer, et al.
Published: (2025)
by: Kloots, Marianne de Heer, et al.
Published: (2025)
Similar Items
-
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
by: Gubian, Michele, et al.
Published: (2025) -
A predictive learning model can simulate temporal dynamics and context effects found in neural representations of continuous speech
by: Liu, Oli Danyi, et al.
Published: (2024) -
The role of audio-visual integration in the time course of phonetic encoding in self-supervised speech models
by: Wang, Yi, et al.
Published: (2025) -
A framework for analyzing concept representations in neural models
by: Naowarat, Burin, et al.
Published: (2026) -
Sustainable self-supervised learning for speech representations
by: Lugo, Luis, et al.
Published: (2024)