Lexicon-Level Contrastive Visual-Grounding Improves Language Modeling

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhuang, Chengxu, Fedorenko, Evelina, Andreas, Jacob
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917619108413440
author Zhuang, Chengxu
Fedorenko, Evelina
Andreas, Jacob
author_facet Zhuang, Chengxu
Fedorenko, Evelina
Andreas, Jacob
contents Today's most accurate language models are trained on orders of magnitude more language data than human language learners receive - but with no supervision from other sensory modalities that play a crucial role in human learning. Can we make LMs' representations and predictions more accurate (and more human-like) with more ecologically plausible supervision? This paper describes LexiContrastive Grounding (LCG), a grounded language learning procedure that leverages visual supervision to improve textual representations. LexiContrastive Grounding combines a next token prediction strategy with a contrastive visual grounding objective, focusing on early-layer representations that encode lexical information. Across multiple word-learning and sentence-understanding benchmarks, LexiContrastive Grounding not only outperforms standard language-only models in learning efficiency, but also improves upon vision-and-language learning procedures including CLIP, GIT, Flamingo, and Vokenization. Moreover, LexiContrastive Grounding improves perplexity by around 5% on multiple language modeling tasks. This work underscores the potential of incorporating visual grounding into language models, aligning more closely with the multimodal nature of human language acquisition.
format Preprint
id arxiv_https___arxiv_org_abs_2403_14551
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Lexicon-Level Contrastive Visual-Grounding Improves Language Modeling
Zhuang, Chengxu
Fedorenko, Evelina
Andreas, Jacob
Computation and Language
Artificial Intelligence
Machine Learning
Today's most accurate language models are trained on orders of magnitude more language data than human language learners receive - but with no supervision from other sensory modalities that play a crucial role in human learning. Can we make LMs' representations and predictions more accurate (and more human-like) with more ecologically plausible supervision? This paper describes LexiContrastive Grounding (LCG), a grounded language learning procedure that leverages visual supervision to improve textual representations. LexiContrastive Grounding combines a next token prediction strategy with a contrastive visual grounding objective, focusing on early-layer representations that encode lexical information. Across multiple word-learning and sentence-understanding benchmarks, LexiContrastive Grounding not only outperforms standard language-only models in learning efficiency, but also improves upon vision-and-language learning procedures including CLIP, GIT, Flamingo, and Vokenization. Moreover, LexiContrastive Grounding improves perplexity by around 5% on multiple language modeling tasks. This work underscores the potential of incorporating visual grounding into language models, aligning more closely with the multimodal nature of human language acquisition.
title Lexicon-Level Contrastive Visual-Grounding Improves Language Modeling
topic Computation and Language
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2403.14551