The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Chen, Xinyi, Fernández, Raquel, Pezzelle, Sandro
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866929215348146176
author Chen, Xinyi
Fernández, Raquel
Pezzelle, Sandro
author_facet Chen, Xinyi
Fernández, Raquel
Pezzelle, Sandro
contents Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what extent they handle basic linguistic constructions -- active-passive voice, coordination, and relative clauses -- that even preschool children can typically master. We present BLA, a novel, automatically constructed benchmark to evaluate multimodal models on these Basic Language Abilities. We show that different types of Transformer-based systems, such as CLIP, ViLBERT, and BLIP2, generally struggle with BLA in a zero-shot setting, in line with previous findings. Our experiments, in particular, show that most of the tested models only marginally benefit when fine-tuned or prompted with construction-specific samples. Yet, the generative BLIP2 shows promising trends, especially in an in-context learning setting. This opens the door to using BLA not only as an evaluation benchmark but also to improve models' basic language abilities.
format Preprint
id arxiv_https___arxiv_org_abs_2310_15061
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
Chen, Xinyi
Fernández, Raquel
Pezzelle, Sandro
Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what extent they handle basic linguistic constructions -- active-passive voice, coordination, and relative clauses -- that even preschool children can typically master. We present BLA, a novel, automatically constructed benchmark to evaluate multimodal models on these Basic Language Abilities. We show that different types of Transformer-based systems, such as CLIP, ViLBERT, and BLIP2, generally struggle with BLA in a zero-shot setting, in line with previous findings. Our experiments, in particular, show that most of the tested models only marginally benefit when fine-tuned or prompted with construction-specific samples. Yet, the generative BLIP2 shows promising trends, especially in an in-context learning setting. This opens the door to using BLA not only as an evaluation benchmark but also to improve models' basic language abilities.
title The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
topic Computation and Language
Artificial Intelligence
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2310.15061