What does Kiki look like? Cross-modal associations between speech sounds and visual shapes in vision-and-language models
Fuente:
arXiv
Saved in:
| Main Authors: | Verhoef, Tessa, Shahrasbi, Kiana, Kouwenhoven, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect
by: Kouwenhoven, Tom, et al.
Published: (2025)
by: Kouwenhoven, Tom, et al.
Published: (2025)
Searching for Structure: Investigating Emergent Communication with Large Language Models
by: Kouwenhoven, Tom, et al.
Published: (2024)
by: Kouwenhoven, Tom, et al.
Published: (2024)
The Curious Case of Representational Alignment: Unravelling Visio-Linguistic Tasks in Emergent Communication
by: Kouwenhoven, Tom, et al.
Published: (2024)
by: Kouwenhoven, Tom, et al.
Published: (2024)
Shaping Shared Languages: Human and Large Language Models' Inductive Biases in Emergent Communication
by: Kouwenhoven, Tom, et al.
Published: (2025)
by: Kouwenhoven, Tom, et al.
Published: (2025)
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
by: Pitta, Elena, et al.
Published: (2025)
by: Pitta, Elena, et al.
Published: (2025)
Memory-Augmented Generative Adversarial Transformers
by: Raaijmakers, Stephan, et al.
Published: (2024)
by: Raaijmakers, Stephan, et al.
Published: (2024)
NeLLCom-X: A Comprehensive Neural-Agent Framework to Simulate Language Learning and Group Communication
by: Lian, Yuchen, et al.
Published: (2024)
by: Lian, Yuchen, et al.
Published: (2024)
Simulating the Emergence of Differential Case Marking with Communicating Neural-Network Agents
by: Lian, Yuchen, et al.
Published: (2025)
by: Lian, Yuchen, et al.
Published: (2025)
Cross-modal linkage risk in clinical vision-language models
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
by: Arasteh, Soroosh Tayebi, et al.
Published: (2026)
What does it take to get state of the art in simultaneous speech-to-speech translation?
by: Wilmet, Vincent, et al.
Published: (2024)
by: Wilmet, Vincent, et al.
Published: (2024)
NeLLCom-Lex: A Neural-agent Framework to Study the Interplay between Lexical Systems and Language Use
by: Zhang, Yuqing, et al.
Published: (2025)
by: Zhang, Yuqing, et al.
Published: (2025)
What does it mean to understand language?
by: Casto, Colton, et al.
Published: (2025)
by: Casto, Colton, et al.
Published: (2025)
Modeling Human-Like Color Naming Behavior in Context
by: Zhang, Yuqing, et al.
Published: (2026)
by: Zhang, Yuqing, et al.
Published: (2026)
The in-context inductive biases of vision-language models differ across modalities
by: Allen, Kelsey, et al.
Published: (2025)
by: Allen, Kelsey, et al.
Published: (2025)
Kiki or Bouba? Sound Symbolism in Vision-and-Language Models
by: Alper, Morris, et al.
Published: (2023)
by: Alper, Morris, et al.
Published: (2023)
Enriching Historical Records: An OCR and AI-Driven Approach for Database Integration
by: Abedi, Zahra, et al.
Published: (2025)
by: Abedi, Zahra, et al.
Published: (2025)
Traces of Social Competence in Large Language Models
by: Kouwenhoven, Tom, et al.
Published: (2026)
by: Kouwenhoven, Tom, et al.
Published: (2026)
Is Temperature the Creativity Parameter of Large Language Models?
by: Peeperkorn, Max, et al.
Published: (2024)
by: Peeperkorn, Max, et al.
Published: (2024)
Mind the Gap: Conformative Decoding to Improve Output Diversity of Instruction-Tuned Large Language Models
by: Peeperkorn, Max, et al.
Published: (2025)
by: Peeperkorn, Max, et al.
Published: (2025)
Analyzing the relationships between pretraining language, phonetic, tonal, and speaker information in self-supervised speech models
by: Gubian, Michele, et al.
Published: (2025)
by: Gubian, Michele, et al.
Published: (2025)
Language-agnostic, automated assessment of listeners' speech recall using large language models
by: Herrmann, Björn
Published: (2025)
by: Herrmann, Björn
Published: (2025)
The mutual exclusivity bias of bilingual visually grounded speech models
by: Oneata, Dan, et al.
Published: (2025)
by: Oneata, Dan, et al.
Published: (2025)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
What does translanguaging look like in an English language classroom in a disadvantaged context? A case in a primary school with difficult circumstances in the Global South
by: Melina Porto
Published: (2025)
by: Melina Porto
Published: (2025)
Context informs pragmatic interpretation in vision-language models
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
by: Tan, Alvin Wei Ming, et al.
Published: (2025)
Transferable speech-to-text large language model alignment module
by: Wu, Boyong, et al.
Published: (2024)
by: Wu, Boyong, et al.
Published: (2024)
Can large audio language models understand child stuttering speech? speech summarization, and source separation
by: Okocha, Chibuzor, et al.
Published: (2025)
by: Okocha, Chibuzor, et al.
Published: (2025)
Protecting multimodal large language models against misleading visualizations
by: Tonglet, Jonathan, et al.
Published: (2025)
by: Tonglet, Jonathan, et al.
Published: (2025)
Diagnosing our datasets: How does my language model learn clinical information?
by: Jia, Furong, et al.
Published: (2025)
by: Jia, Furong, et al.
Published: (2025)
Improving child speech recognition with augmented child-like speech
by: Zhang, Yuanyuan, et al.
Published: (2024)
by: Zhang, Yuanyuan, et al.
Published: (2024)
A closer look at how large language models trust humans: patterns and biases
by: Lerman, Valeria, et al.
Published: (2025)
by: Lerman, Valeria, et al.
Published: (2025)
What do self-supervised speech models know about Dutch? Analyzing advantages of language-specific pre-training
by: Kloots, Marianne de Heer, et al.
Published: (2025)
by: Kloots, Marianne de Heer, et al.
Published: (2025)
How does fine-tuning improve sensorimotor representations in large language models?
by: Wu, Minghua, et al.
Published: (2026)
by: Wu, Minghua, et al.
Published: (2026)
Automating construction safety inspections using a multi-modal vision-language RAG framework
by: Wang, Chenxin, et al.
Published: (2025)
by: Wang, Chenxin, et al.
Published: (2025)
Verifying Cross-modal Entity Consistency in News using Vision-language Models
by: Tahmasebi, Sahar, et al.
Published: (2025)
by: Tahmasebi, Sahar, et al.
Published: (2025)
Investigating large language models for their competence in extracting grammatically sound sentences from transcribed noisy utterances
by: Wróblewska, Alina
Published: (2024)
by: Wróblewska, Alina
Published: (2024)
Dual language intervention in a case of severe speech sound disorde
by: Eliane Ramos
Published: (2014)
by: Eliane Ramos
Published: (2014)
MNAFT: modality neuron-aware fine-tuning of multimodal large language models for image translation
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
Ensemble of pre-trained language models and data augmentation for hate speech detection from Arabic tweets
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
by: Daouadi, Kheir Eddine, et al.
Published: (2024)
AugSumm: towards generalizable speech summarization using synthetic labels from large language model
by: Jung, Jee-weon, et al.
Published: (2024)
by: Jung, Jee-weon, et al.
Published: (2024)
Similar Items
-
Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki Effect
by: Kouwenhoven, Tom, et al.
Published: (2025) -
Searching for Structure: Investigating Emergent Communication with Large Language Models
by: Kouwenhoven, Tom, et al.
Published: (2024) -
The Curious Case of Representational Alignment: Unravelling Visio-Linguistic Tasks in Emergent Communication
by: Kouwenhoven, Tom, et al.
Published: (2024) -
Shaping Shared Languages: Human and Large Language Models' Inductive Biases in Emergent Communication
by: Kouwenhoven, Tom, et al.
Published: (2025) -
Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
by: Pitta, Elena, et al.
Published: (2025)