Guardado en:
| Autores principales: | van Rensburg, Kyle Janse, van Niekerk, Benjamin, Kamper, Herman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2603.03096 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Revisiting speech segmentation and lexicon learning with better features
por: Kamper, Herman, et al.
Publicado: (2024)
por: Kamper, Herman, et al.
Publicado: (2024)
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
por: Malan, Simon, et al.
Publicado: (2024)
por: Malan, Simon, et al.
Publicado: (2024)
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
por: Malan, Simon, et al.
Publicado: (2025)
por: Malan, Simon, et al.
Publicado: (2025)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
por: Kamper, Herman, et al.
Publicado: (2025)
por: Kamper, Herman, et al.
Publicado: (2025)
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
por: Visser, Nicol, et al.
Publicado: (2025)
por: Visser, Nicol, et al.
Publicado: (2025)
Spoken-Term Discovery using Discrete Speech Units
por: van Niekerk, Benjamin, et al.
Publicado: (2024)
por: van Niekerk, Benjamin, et al.
Publicado: (2024)
Disentanglement in a GAN for Unconditional Speech Synthesis
por: Baas, Matthew, et al.
Publicado: (2023)
por: Baas, Matthew, et al.
Publicado: (2023)
Visually Grounded Speech Models have a Mutual Exclusivity Bias
por: Nortje, Leanne, et al.
Publicado: (2024)
por: Nortje, Leanne, et al.
Publicado: (2024)
Translating speech with just images
por: Oneata, Dan, et al.
Publicado: (2024)
por: Oneata, Dan, et al.
Publicado: (2024)
Visually grounded few-shot word learning in low-resource settings
por: Nortje, Leanne, et al.
Publicado: (2023)
por: Nortje, Leanne, et al.
Publicado: (2023)
Towards few-shot isolated word reading assessment
por: Smit, Reuben, et al.
Publicado: (2025)
por: Smit, Reuben, et al.
Publicado: (2025)
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
por: Baas, Matthew, et al.
Publicado: (2025)
por: Baas, Matthew, et al.
Publicado: (2025)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
por: Slabbert, Danel, et al.
Publicado: (2025)
por: Slabbert, Danel, et al.
Publicado: (2025)
Feature-based analysis of oral narratives from Afrikaans and isiXhosa children
por: Sharratt, Emma, et al.
Publicado: (2025)
por: Sharratt, Emma, et al.
Publicado: (2025)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
por: Visser, Nicol, et al.
Publicado: (2026)
por: Visser, Nicol, et al.
Publicado: (2026)
The mutual exclusivity bias of bilingual visually grounded speech models
por: Oneata, Dan, et al.
Publicado: (2025)
por: Oneata, Dan, et al.
Publicado: (2025)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
por: Jacobs, Christiaan, et al.
Publicado: (2025)
por: Jacobs, Christiaan, et al.
Publicado: (2025)
Identifying Speaker Information in Feed-Forward Layers of Self-Supervised Speech Transformers
por: Lin, Tzu-Quan, et al.
Publicado: (2025)
por: Lin, Tzu-Quan, et al.
Publicado: (2025)
Linear-Complexity Self-Supervised Learning for Speech Processing
por: Zhang, Shucong, et al.
Publicado: (2024)
por: Zhang, Shucong, et al.
Publicado: (2024)
ELF: Encoding Speaker-Specific Latent Speech Feature for Speech Synthesis
por: Kong, Jungil, et al.
Publicado: (2023)
por: Kong, Jungil, et al.
Publicado: (2023)
Automatically assessing oral narratives of Afrikaans and isiXhosa children
por: Louw, Retief, et al.
Publicado: (2025)
por: Louw, Retief, et al.
Publicado: (2025)
DisfluencySpeech -- Single-Speaker Conversational Speech Dataset with Paralanguage
por: Wang, Kyra, et al.
Publicado: (2024)
por: Wang, Kyra, et al.
Publicado: (2024)
Investigation of Speaker Representation for Target-Speaker Speech Processing
por: Ashihara, Takanori, et al.
Publicado: (2024)
por: Ashihara, Takanori, et al.
Publicado: (2024)
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
por: Nortje, Leanne, et al.
Publicado: (2024)
por: Nortje, Leanne, et al.
Publicado: (2024)
Textless Acoustic Model with Self-Supervised Distillation for Noise-Robust Expressive Speech-to-Speech Translation
por: Hwang, Min-Jae, et al.
Publicado: (2024)
por: Hwang, Min-Jae, et al.
Publicado: (2024)
Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
por: Komatsu, Ryota, et al.
Publicado: (2024)
por: Komatsu, Ryota, et al.
Publicado: (2024)
What Do Self-Supervised Speech and Speaker Models Learn? New Findings From a Cross Model Layer-Wise Analysis
por: Ashihara, Takanori, et al.
Publicado: (2024)
por: Ashihara, Takanori, et al.
Publicado: (2024)
Self-Supervised Models of Speech Infer Universal Articulatory Kinematics
por: Cho, Cheol Jun, et al.
Publicado: (2023)
por: Cho, Cheol Jun, et al.
Publicado: (2023)
Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs
por: Tseng, Wei-Cheng, et al.
Publicado: (2025)
por: Tseng, Wei-Cheng, et al.
Publicado: (2025)
Robust Unsupervised Adaptation of a Speech Recogniser Using Entropy Minimisation and Speaker Codes
por: van Dalen, Rogier C., et al.
Publicado: (2025)
por: van Dalen, Rogier C., et al.
Publicado: (2025)
Interface Design for Self-Supervised Speech Models
por: Shih, Yi-Jen, et al.
Publicado: (2024)
por: Shih, Yi-Jen, et al.
Publicado: (2024)
Optimizing Speech-Input Length for Speaker-Independent Depression Classification
por: Rutowski, Tomasz, et al.
Publicado: (2024)
por: Rutowski, Tomasz, et al.
Publicado: (2024)
Exploring Effective Distillation of Self-Supervised Speech Models for Automatic Speech Recognition
por: Wang, Yujin, et al.
Publicado: (2022)
por: Wang, Yujin, et al.
Publicado: (2022)
Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
por: Blandón, María Andrea Cruz, et al.
Publicado: (2025)
por: Blandón, María Andrea Cruz, et al.
Publicado: (2025)
Speaker-Distinguishable CTC: Learning Speaker Distinction Using CTC for Multi-Talker Speech Recognition
por: Sakuma, Asahi, et al.
Publicado: (2025)
por: Sakuma, Asahi, et al.
Publicado: (2025)
Speaker-Aware Simulation Improves Conversational Speech Recognition
por: Gedeon, Máté, et al.
Publicado: (2026)
por: Gedeon, Máté, et al.
Publicado: (2026)
DiariST: Streaming Speech Translation with Speaker Diarization
por: Yang, Mu, et al.
Publicado: (2023)
por: Yang, Mu, et al.
Publicado: (2023)
TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning Models
por: Peng, Junyi, et al.
Publicado: (2025)
por: Peng, Junyi, et al.
Publicado: (2025)
STaR: Distilling Speech Temporal Relation for Lightweight Speech Self-Supervised Learning Models
por: Jang, Kangwook, et al.
Publicado: (2023)
por: Jang, Kangwook, et al.
Publicado: (2023)
Ejemplares similares
-
Revisiting speech segmentation and lexicon learning with better features
por: Kamper, Herman, et al.
Publicado: (2024) -
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
por: Malan, Simon, et al.
Publicado: (2024) -
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
por: Malan, Simon, et al.
Publicado: (2025) -
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
por: Carbonneau, Marc-André, et al.
Publicado: (2025) -
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
por: Kamper, Herman, et al.
Publicado: (2025)