The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Sauter, Adrian, Zuidema, Willem, Kloots, Marianne de Heer
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911163482112000
author Sauter, Adrian
Zuidema, Willem
Kloots, Marianne de Heer
author_facet Sauter, Adrian
Zuidema, Willem
Kloots, Marianne de Heer
contents How does visual information included in training affect language processing in audio- and text-based deep learning models? We explore how such visual grounding affects model-internal representations of words, and find substantially different effects in speech- vs. text-based language encoders. Firstly, global representational comparisons reveal that visual grounding increases alignment between representations of spoken and written language, but this effect seems mainly driven by enhanced encoding of word identity rather than meaning. We then apply targeted clustering analyses to probe for phonetic vs. semantic discriminability in model representations. Speech-based representations remain phonetically dominated with visual grounding, but in contrast to text-based representations, visual grounding does not improve semantic discriminability. Our findings could usefully inform the development of more efficient methods to enrich speech-based models with visually-informed semantics.
format Preprint
id arxiv_https___arxiv_org_abs_2509_15837
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
Sauter, Adrian
Zuidema, Willem
Kloots, Marianne de Heer
Computation and Language
I.2.7
How does visual information included in training affect language processing in audio- and text-based deep learning models? We explore how such visual grounding affects model-internal representations of words, and find substantially different effects in speech- vs. text-based language encoders. Firstly, global representational comparisons reveal that visual grounding increases alignment between representations of spoken and written language, but this effect seems mainly driven by enhanced encoding of word identity rather than meaning. We then apply targeted clustering analyses to probe for phonetic vs. semantic discriminability in model representations. Speech-based representations remain phonetically dominated with visual grounding, but in contrast to text-based representations, visual grounding does not improve semantic discriminability. Our findings could usefully inform the development of more efficient methods to enrich speech-based models with visually-informed semantics.
title The Curious Case of Visual Grounding: Different Effects for Speech- and Text-based Language Encoders
topic Computation and Language
I.2.7
url https://arxiv.org/abs/2509.15837