Kiki or Bouba? Sound Symbolism in Vision-and-Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Alper, Morris, Averbuch-Elor, Hadar
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913294452785152
author Alper, Morris
Averbuch-Elor, Hadar
author_facet Alper, Morris
Averbuch-Elor, Hadar
contents Although the mapping between sound and meaning in human language is assumed to be largely arbitrary, research in cognitive science has shown that there are non-trivial correlations between particular sounds and meanings across languages and demographic groups, a phenomenon known as sound symbolism. Among the many dimensions of meaning, sound symbolism is particularly salient and well-demonstrated with regards to cross-modal associations between language and the visual domain. In this work, we address the question of whether sound symbolism is reflected in vision-and-language models such as CLIP and Stable Diffusion. Using zero-shot knowledge probing to investigate the inherent knowledge of these models, we find strong evidence that they do show this pattern, paralleling the well-known kiki-bouba effect in psycholinguistics. Our work provides a novel method for demonstrating sound symbolism and understanding its nature using computational tools. Our code will be made publicly available.
format Preprint
id arxiv_https___arxiv_org_abs_2310_16781
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Kiki or Bouba? Sound Symbolism in Vision-and-Language Models
Alper, Morris
Averbuch-Elor, Hadar
Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
Although the mapping between sound and meaning in human language is assumed to be largely arbitrary, research in cognitive science has shown that there are non-trivial correlations between particular sounds and meanings across languages and demographic groups, a phenomenon known as sound symbolism. Among the many dimensions of meaning, sound symbolism is particularly salient and well-demonstrated with regards to cross-modal associations between language and the visual domain. In this work, we address the question of whether sound symbolism is reflected in vision-and-language models such as CLIP and Stable Diffusion. Using zero-shot knowledge probing to investigate the inherent knowledge of these models, we find strong evidence that they do show this pattern, paralleling the well-known kiki-bouba effect in psycholinguistics. Our work provides a novel method for demonstrating sound symbolism and understanding its nature using computational tools. Our code will be made publicly available.
title Kiki or Bouba? Sound Symbolism in Vision-and-Language Models
topic Computer Vision and Pattern Recognition
Computation and Language
Machine Learning
url https://arxiv.org/abs/2310.16781