Visually Grounded Speech Models have a Mutual Exclusivity Bias
Fuente:
arXiv
Saved in:
| Main Authors: | Nortje, Leanne, Oneaţă, Dan, Matusevych, Yevgen, Kamper, Herman |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The mutual exclusivity bias of bilingual visually grounded speech models
by: Oneata, Dan, et al.
Published: (2025)
by: Oneata, Dan, et al.
Published: (2025)
Visually grounded few-shot word learning in low-resource settings
by: Nortje, Leanne, et al.
Published: (2023)
by: Nortje, Leanne, et al.
Published: (2023)
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
by: Nortje, Leanne, et al.
Published: (2024)
by: Nortje, Leanne, et al.
Published: (2024)
Translating speech with just images
by: Oneata, Dan, et al.
Published: (2024)
by: Oneata, Dan, et al.
Published: (2024)
Disentanglement in a GAN for Unconditional Speech Synthesis
by: Baas, Matthew, et al.
Published: (2023)
by: Baas, Matthew, et al.
Published: (2023)
Spoken Language Modeling with Duration-Penalized Self-Supervised Units
by: Visser, Nicol, et al.
Published: (2025)
by: Visser, Nicol, et al.
Published: (2025)
Interpreting Speaker Characteristics in the Dimensions of Self-Supervised Speech Features
by: van Rensburg, Kyle Janse, et al.
Published: (2026)
by: van Rensburg, Kyle Janse, et al.
Published: (2026)
MARS6: A Small and Robust Hierarchical-Codec Text-to-Speech Model
by: Baas, Matthew, et al.
Published: (2025)
by: Baas, Matthew, et al.
Published: (2025)
Towards few-shot isolated word reading assessment
by: Smit, Reuben, et al.
Published: (2025)
by: Smit, Reuben, et al.
Published: (2025)
ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling
by: Visser, Nicol, et al.
Published: (2026)
by: Visser, Nicol, et al.
Published: (2026)
Revisiting speech segmentation and lexicon learning with better features
by: Kamper, Herman, et al.
Published: (2024)
by: Kamper, Herman, et al.
Published: (2024)
Unsupervised lexicon learning from speech is limited by representations rather than clustering
by: Slabbert, Danel, et al.
Published: (2025)
by: Slabbert, Danel, et al.
Published: (2025)
Unsupervised Word Discovery: Boundary Detection with Clustering vs. Dynamic Programming
by: Malan, Simon, et al.
Published: (2024)
by: Malan, Simon, et al.
Published: (2024)
Should Top-Down Clustering Affect Boundaries in Unsupervised Word Discovery?
by: Malan, Simon, et al.
Published: (2025)
by: Malan, Simon, et al.
Published: (2025)
Speech Recognition for Automatically Assessing Afrikaans and isiXhosa Preschool Oral Narratives
by: Jacobs, Christiaan, et al.
Published: (2025)
by: Jacobs, Christiaan, et al.
Published: (2025)
LinearVC: Linear transformations of self-supervised features through the lens of voice conversion
by: Kamper, Herman, et al.
Published: (2025)
by: Kamper, Herman, et al.
Published: (2025)
Automatically assessing oral narratives of Afrikaans and isiXhosa children
by: Louw, Retief, et al.
Published: (2025)
by: Louw, Retief, et al.
Published: (2025)
Feature-based analysis of oral narratives from Afrikaans and isiXhosa children
by: Sharratt, Emma, et al.
Published: (2025)
by: Sharratt, Emma, et al.
Published: (2025)
Analyzing and Improving Speaker Similarity Assessment for Speech Synthesis
by: Carbonneau, Marc-André, et al.
Published: (2025)
by: Carbonneau, Marc-André, et al.
Published: (2025)
Spoken Stereoset: On Evaluating Social Bias Toward Speaker in Speech Large Language Models
by: Lin, Yi-Cheng, et al.
Published: (2024)
by: Lin, Yi-Cheng, et al.
Published: (2024)
Spoken-Term Discovery using Discrete Speech Units
by: van Niekerk, Benjamin, et al.
Published: (2024)
by: van Niekerk, Benjamin, et al.
Published: (2024)
Revisiting Self-supervised Learning of Speech Representation from a Mutual Information Perspective
by: Liu, Alexander H., et al.
Published: (2024)
by: Liu, Alexander H., et al.
Published: (2024)
Leveraging Audio-Visual Data to Reduce the Multilingual Gap in Self-Supervised Speech Models
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
by: Blandón, María Andrea Cruz, et al.
Published: (2025)
Speak Your Mind: The Speech Continuation Task as a Probe of Voice-Based Model Bias
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2025)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2025)
TagSpeech: End-to-End Multi-Speaker ASR and Diarization with Fine-Grained Temporal Grounding
by: Huo, Mingyue, et al.
Published: (2026)
by: Huo, Mingyue, et al.
Published: (2026)
Gender Bias in Instruction-Guided Speech Synthesis Models
by: Kuan, Chun-Yi, et al.
Published: (2025)
by: Kuan, Chun-Yi, et al.
Published: (2025)
Towards generalisable and calibrated synthetic speech detection with self-supervised representations
by: Pascu, Octavian, et al.
Published: (2023)
by: Pascu, Octavian, et al.
Published: (2023)
The Voice Behind the Words: Quantifying Intersectional Bias in SpeechLLMs
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
by: Satish, Shree Harsha Bokkahalli, et al.
Published: (2026)
A Contrastive Learning Approach to Mitigate Bias in Speech Models
by: Koudounas, Alkis, et al.
Published: (2024)
by: Koudounas, Alkis, et al.
Published: (2024)
Visual-Aware Speech Recognition for Noisy Scenarios
by: Balaji, Lakshmipathi, et al.
Published: (2025)
by: Balaji, Lakshmipathi, et al.
Published: (2025)
How to Evaluate Automatic Speech Recognition: Comparing Different Performance and Bias Measures
by: Patel, Tanvina, et al.
Published: (2025)
by: Patel, Tanvina, et al.
Published: (2025)
DeSTA: Enhancing Speech Language Models through Descriptive Speech-Text Alignment
by: Lu, Ke-Han, et al.
Published: (2024)
by: Lu, Ke-Han, et al.
Published: (2024)
Towards Efficient Speech-Text Jointly Decoding within One Speech Language Model
by: Wu, Haibin, et al.
Published: (2025)
by: Wu, Haibin, et al.
Published: (2025)
Speech Recognition Model Improves Text-to-Speech Synthesis using Fine-Grained Reward
by: Wang, Guansu, et al.
Published: (2025)
by: Wang, Guansu, et al.
Published: (2025)
Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer
by: Zhu, Yongxin, et al.
Published: (2024)
by: Zhu, Yongxin, et al.
Published: (2024)
Speech-IFEval: Evaluating Instruction-Following and Quantifying Catastrophic Forgetting in Speech-Aware Language Models
by: Lu, Ke-Han, et al.
Published: (2025)
by: Lu, Ke-Han, et al.
Published: (2025)
VIBE: Voice-Induced open-ended Bias Evaluation for Large Audio-Language Models via Real-World Speech
by: Lin, Yi-Cheng, et al.
Published: (2026)
by: Lin, Yi-Cheng, et al.
Published: (2026)
Contextualized Automatic Speech Recognition with Attention-Based Bias Phrase Boosted Beam Search
by: Sudo, Yui, et al.
Published: (2024)
by: Sudo, Yui, et al.
Published: (2024)
Listen and Speak Fairly: A Study on Semantic Gender Bias in Speech Integrated Large Language Models
by: Lin, Yi-Cheng, et al.
Published: (2024)
by: Lin, Yi-Cheng, et al.
Published: (2024)
SpeechCaps: Advancing Instruction-Based Universal Speech Models with Multi-Talker Speaking Style Captioning
by: Huang, Chien-yu, et al.
Published: (2024)
by: Huang, Chien-yu, et al.
Published: (2024)
Similar Items
-
The mutual exclusivity bias of bilingual visually grounded speech models
by: Oneata, Dan, et al.
Published: (2025) -
Visually grounded few-shot word learning in low-resource settings
by: Nortje, Leanne, et al.
Published: (2023) -
Improved Visually Prompted Keyword Localisation in Real Low-Resource Settings
by: Nortje, Leanne, et al.
Published: (2024) -
Translating speech with just images
by: Oneata, Dan, et al.
Published: (2024) -
Disentanglement in a GAN for Unconditional Speech Synthesis
by: Baas, Matthew, et al.
Published: (2023)