The mutual exclusivity bias of bilingual visually grounded speech models
Fuente:
arXiv
Salvato in:
| Autori principali: | Oneata, Dan, Nortje, Leanne, Matusevych, Yevgen, Kamper, Herman |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Visually Grounded Speech Models have a Mutual Exclusivity Bias
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
Visually grounded few-shot word learning in low-resource settings
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
di: Nortje, Leanne, et al.
Pubblicazione: (2023)
Translating speech with just images
di: Oneata, Dan, et al.
Pubblicazione: (2024)
di: Oneata, Dan, et al.
Pubblicazione: (2024)
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
di: Nortje, Leanne, et al.
Pubblicazione: (2024)
Revisiting speech segmentation and lexicon learning with better features
di: Kamper, Herman, et al.
Pubblicazione: (2024)
di: Kamper, Herman, et al.
Pubblicazione: (2024)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
di: Slabbert, Danel, et al.
Pubblicazione: (2025)
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
di: Visser, Nicol, et al.
Pubblicazione: (2025)
di: Visser, Nicol, et al.
Pubblicazione: (2025)
Disentanglement in a GAN for Unconditional Speech Synthesis
di: Baas, Matthew, et al.
Pubblicazione: (2023)
di: Baas, Matthew, et al.
Pubblicazione: (2023)
Towards few-shot isolated word reading assessment
di: Smit, Reuben, et al.
Pubblicazione: (2025)
di: Smit, Reuben, et al.
Pubblicazione: (2025)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
di: Visser, Nicol, et al.
Pubblicazione: (2026)
di: Visser, Nicol, et al.
Pubblicazione: (2026)
Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
di: van Rensburg, Kyle Janse, et al.
Pubblicazione: (2026)
di: van Rensburg, Kyle Janse, et al.
Pubblicazione: (2026)
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
di: Malan, Simon, et al.
Pubblicazione: (2025)
di: Malan, Simon, et al.
Pubblicazione: (2025)
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
di: Malan, Simon, et al.
Pubblicazione: (2024)
di: Malan, Simon, et al.
Pubblicazione: (2024)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
di: Pascu, Octavian, et al.
Pubblicazione: (2023)
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
di: Kamper, Herman, et al.
Pubblicazione: (2025)
di: Kamper, Herman, et al.
Pubblicazione: (2025)
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
di: Baas, Matthew, et al.
Pubblicazione: (2025)
di: Baas, Matthew, et al.
Pubblicazione: (2025)
Automatically assessing oral narratives of Afrikaans and isiXhosa children
di: Louw, Retief, et al.
Pubblicazione: (2025)
di: Louw, Retief, et al.
Pubblicazione: (2025)
Feature-based analysis of oral narratives from Afrikaans and isiXhosa children
di: Sharratt, Emma, et al.
Pubblicazione: (2025)
di: Sharratt, Emma, et al.
Pubblicazione: (2025)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
di: Jacobs, Christiaan, et al.
Pubblicazione: (2025)
di: Jacobs, Christiaan, et al.
Pubblicazione: (2025)
emg2speech: Synthesizing speech from electromyography using self-supervised speech models
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
di: Gowda, Harshavardhana T., et al.
Pubblicazione: (2025)
Prominence-aware automatic speech recognition for conversational speech
di: Linke, Julian, et al.
Pubblicazione: (2025)
di: Linke, Julian, et al.
Pubblicazione: (2025)
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
di: Lyth, Dan, et al.
Pubblicazione: (2024)
di: Lyth, Dan, et al.
Pubblicazione: (2024)
Predicting positive transfer for improved low-resource speech recognition using acoustic pseudo-tokens
di: San, Nay, et al.
Pubblicazione: (2024)
di: San, Nay, et al.
Pubblicazione: (2024)
XCB: an effective contextual biasing approach to bias cross-lingual phrases in speech recognition
di: Wan, Xucheng, et al.
Pubblicazione: (2024)
di: Wan, Xucheng, et al.
Pubblicazione: (2024)
WavLM model ensemble for audio deepfake detection
di: Combei, David, et al.
Pubblicazione: (2024)
di: Combei, David, et al.
Pubblicazione: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
di: Okocha, Chibuzor, et al.
Pubblicazione: (2025)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
di: Gubian, Michele, et al.
Pubblicazione: (2025)
di: Gubian, Michele, et al.
Pubblicazione: (2025)
On the reliability of feature attribution methods for speech classification
di: Shen, Gaofei, et al.
Pubblicazione: (2025)
di: Shen, Gaofei, et al.
Pubblicazione: (2025)
Circumventing shortcuts in audio-visual deepfake detection datasets with unsupervised learning
di: Smeu, Stefan, et al.
Pubblicazione: (2024)
di: Smeu, Stefan, et al.
Pubblicazione: (2024)
Improving child speech recognition with augmented child-like speech
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yuanyuan, et al.
Pubblicazione: (2024)
Transferable speech-to-text large language model alignment module
di: Wu, Boyong, et al.
Pubblicazione: (2024)
di: Wu, Boyong, et al.
Pubblicazione: (2024)
Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2025)
di: Kashiwagi, Yosuke, et al.
Pubblicazione: (2025)
Disentangling segmental and prosodic factors to non-native speech comprehensibility
di: Quamer, Waris, et al.
Pubblicazione: (2024)
di: Quamer, Waris, et al.
Pubblicazione: (2024)
Exploring the limits of decoder-only models trained on public speech recognition corpora
di: Gupta, Ankit, et al.
Pubblicazione: (2024)
di: Gupta, Ankit, et al.
Pubblicazione: (2024)
IIITH-BUT system for IWSLT 2025 low-resource Bhojpuri to Hindi speech translation
di: Akkiraju, Bhavana, et al.
Pubblicazione: (2025)
di: Akkiraju, Bhavana, et al.
Pubblicazione: (2025)
Can we reconstruct a dysarthric voice with the large speech model Parler TTS?
di: Sanchez, Ariadna, et al.
Pubblicazione: (2025)
di: Sanchez, Ariadna, et al.
Pubblicazione: (2025)
AfriHuBERT: A self-supervised speech representation model for African languages
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
di: Alabi, Jesujoba O., et al.
Pubblicazione: (2024)
1000 African Voices: Advancing inclusive multi-speaker multi-accent speech synthesis
di: Ogun, Sewade, et al.
Pubblicazione: (2024)
di: Ogun, Sewade, et al.
Pubblicazione: (2024)
SpeechCLIP+: Self-supervised multi-task representation learning for speech via CLIP and speech-image data
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
di: Wang, Hsuan-Fu, et al.
Pubblicazione: (2024)
Training dynamic models using early exits for automatic speech recognition on resource-constrained devices
di: Wright, George August, et al.
Pubblicazione: (2023)
di: Wright, George August, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Visually Grounded Speech Models have a Mutual Exclusivity Bias
di: Nortje, Leanne, et al.
Pubblicazione: (2024) -
Visually grounded few-shot word learning in low-resource settings
di: Nortje, Leanne, et al.
Pubblicazione: (2023) -
Translating speech with just images
di: Oneata, Dan, et al.
Pubblicazione: (2024) -
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
di: Nortje, Leanne, et al.
Pubblicazione: (2024) -
Revisiting speech segmentation and lexicon learning with better features
di: Kamper, Herman, et al.
Pubblicazione: (2024)