Visually grounded few-shot word learning in low-resource settings
Fuente:
arXiv
Guardado en:
| Autores principales: | Nortje, Leanne, Oneata, Dan, Kamper, Herman |
|---|---|
| Formato: | Preprint |
| Publicado: |
2023
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The mutual exclusivity bias of bilingual visually grounded speech models
por: Oneata, Dan, et al.
Publicado: (2025)
por: Oneata, Dan, et al.
Publicado: (2025)
Visually Grounded Speech Models have a Mutual Exclusivity Bias
por: Nortje, Leanne, et al.
Publicado: (2024)
por: Nortje, Leanne, et al.
Publicado: (2024)
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
por: Nortje, Leanne, et al.
Publicado: (2024)
por: Nortje, Leanne, et al.
Publicado: (2024)
Towards few-shot isolated word reading assessment
por: Smit, Reuben, et al.
Publicado: (2025)
por: Smit, Reuben, et al.
Publicado: (2025)
Translating speech with just images
por: Oneata, Dan, et al.
Publicado: (2024)
por: Oneata, Dan, et al.
Publicado: (2024)
Revisiting speech segmentation and lexicon learning with better features
por: Kamper, Herman, et al.
Publicado: (2024)
por: Kamper, Herman, et al.
Publicado: (2024)
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
por: Visser, Nicol, et al.
Publicado: (2025)
por: Visser, Nicol, et al.
Publicado: (2025)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
por: Slabbert, Danel, et al.
Publicado: (2025)
por: Slabbert, Danel, et al.
Publicado: (2025)
Disentanglement in a GAN for Unconditional Speech Synthesis
por: Baas, Matthew, et al.
Publicado: (2023)
por: Baas, Matthew, et al.
Publicado: (2023)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
por: Visser, Nicol, et al.
Publicado: (2026)
por: Visser, Nicol, et al.
Publicado: (2026)
Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
por: van Rensburg, Kyle Janse, et al.
Publicado: (2026)
por: van Rensburg, Kyle Janse, et al.
Publicado: (2026)
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
por: Malan, Simon, et al.
Publicado: (2024)
por: Malan, Simon, et al.
Publicado: (2024)
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
por: Malan, Simon, et al.
Publicado: (2025)
por: Malan, Simon, et al.
Publicado: (2025)
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
por: Kamper, Herman, et al.
Publicado: (2025)
por: Kamper, Herman, et al.
Publicado: (2025)
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
por: Baas, Matthew, et al.
Publicado: (2025)
por: Baas, Matthew, et al.
Publicado: (2025)
Automatically assessing oral narratives of Afrikaans and isiXhosa children
por: Louw, Retief, et al.
Publicado: (2025)
por: Louw, Retief, et al.
Publicado: (2025)
Feature-based analysis of oral narratives from Afrikaans and isiXhosa children
por: Sharratt, Emma, et al.
Publicado: (2025)
por: Sharratt, Emma, et al.
Publicado: (2025)
Multilingual acoustic word embeddings for zero-resource languages
por: Jacobs, Christiaan
Publicado: (2024)
por: Jacobs, Christiaan
Publicado: (2024)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
por: Jacobs, Christiaan, et al.
Publicado: (2025)
por: Jacobs, Christiaan, et al.
Publicado: (2025)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
por: San, Nay, et al.
Publicado: (2024)
por: San, Nay, et al.
Publicado: (2024)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
por: Akkiraju, Bhavana, et al.
Publicado: (2025)
por: Akkiraju, Bhavana, et al.
Publicado: (2025)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
por: Pascu, Octavian, et al.
Publicado: (2023)
por: Pascu, Octavian, et al.
Publicado: (2023)
Target word activity detector: An approach to obtain ASR word boundaries without lexicon
por: Sivasankaran, Sunit, et al.
Publicado: (2024)
por: Sivasankaran, Sunit, et al.
Publicado: (2024)
An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
por: Gong, Cheng, et al.
Publicado: (2024)
por: Gong, Cheng, et al.
Publicado: (2024)
Selective Attention Merging for low resource tasks: A case study of Child ASR
por: Shankar, Natarajan Balaji, et al.
Publicado: (2025)
por: Shankar, Natarajan Balaji, et al.
Publicado: (2025)
Strategies for improving low resource speech to text translation relying on pre-trained ASR models
por: Kesiraju, Santosh, et al.
Publicado: (2023)
por: Kesiraju, Santosh, et al.
Publicado: (2023)
Custom Data Augmentation for low resource ASR using Bark and Retrieval-Based Voice Conversion
por: Kamble, Anand, et al.
Publicado: (2023)
por: Kamble, Anand, et al.
Publicado: (2023)
The Greek podcast corpus: Competitive speech models for low-resourced languages with weakly supervised data
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
por: Paraskevopoulos, Georgios, et al.
Publicado: (2024)
WavLM model ensemble for audio deepfake detection
por: Combei, David, et al.
Publicado: (2024)
por: Combei, David, et al.
Publicado: (2024)
TADA: Training-free Attribution and Out-of-Domain Detection of Audio Deepfakes
por: Stan, Adriana, et al.
Publicado: (2025)
por: Stan, Adriana, et al.
Publicado: (2025)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
por: Carbonneau, Marc-André, et al.
Publicado: (2025)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
por: Smeu, Stefan, et al.
Publicado: (2024)
por: Smeu, Stefan, et al.
Publicado: (2024)
Zero-resource Speech Translation and Recognition with LLMs
por: Mundnich, Karel, et al.
Publicado: (2024)
por: Mundnich, Karel, et al.
Publicado: (2024)
Understanding Zero-shot Rare Word Recognition Improvements Through LLM Integration
por: Wang, Haoxuan
Publicado: (2025)
por: Wang, Haoxuan
Publicado: (2025)
Zero-shot Context Biasing with Trie-based Decoding using Synthetic Multi-Pronunciation
por: Liu, Changsong, et al.
Publicado: (2025)
por: Liu, Changsong, et al.
Publicado: (2025)
Improving noisy student training for low-resource languages in End-to-End ASR using CycleGAN and inter-domain losses
por: Li, Chia-Yu, et al.
Publicado: (2024)
por: Li, Chia-Yu, et al.
Publicado: (2024)
Spoken-Term Discovery using Discrete Speech Units
por: van Niekerk, Benjamin, et al.
Publicado: (2024)
por: van Niekerk, Benjamin, et al.
Publicado: (2024)
A corpus-based investigation of pitch contours of monosyllabic words in conversational Taiwan Mandarin
por: Jin, Xiaoyun, et al.
Publicado: (2024)
por: Jin, Xiaoyun, et al.
Publicado: (2024)
ELAICHI: Enhancing Low-resource TTS by Addressing Infrequent and Low-frequency Character Bigrams
por: Anand, Srija, et al.
Publicado: (2024)
por: Anand, Srija, et al.
Publicado: (2024)
Unmasking real-world audio deepfakes: A data-centric approach
por: Combei, David, et al.
Publicado: (2025)
por: Combei, David, et al.
Publicado: (2025)
Ejemplares similares
-
The mutual exclusivity bias of bilingual visually grounded speech models
por: Oneata, Dan, et al.
Publicado: (2025) -
Visually Grounded Speech Models have a Mutual Exclusivity Bias
por: Nortje, Leanne, et al.
Publicado: (2024) -
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
por: Nortje, Leanne, et al.
Publicado: (2024) -
Towards few-shot isolated word reading assessment
por: Smit, Reuben, et al.
Publicado: (2025) -
Translating speech with just images
por: Oneata, Dan, et al.
Publicado: (2024)