The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models
Fuente:
arXiv
Saved in:
| Main Authors: | , , |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
| _version_ | 1866929215348146176 |
|---|---|
| author | Chen, Xinyi Fernández, Raquel Pezzelle, Sandro |
| author_facet | Chen, Xinyi Fernández, Raquel Pezzelle, Sandro |
| contents | Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what extent they handle basic linguistic constructions -- active-passive voice, coordination, and relative clauses -- that even preschool children can typically master. We present BLA, a novel, automatically constructed benchmark to evaluate multimodal models on these Basic Language Abilities. We show that different types of Transformer-based systems, such as CLIP, ViLBERT, and BLIP2, generally struggle with BLA in a zero-shot setting, in line with previous findings. Our experiments, in particular, show that most of the tested models only marginally benefit when fine-tuned or prompted with construction-specific samples. Yet, the generative BLIP2 shows promising trends, especially in an in-context learning setting. This opens the door to using BLA not only as an evaluation benchmark but also to improve models' basic language abilities. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2310_15061 |
| institution | arXiv |
| publishDate | 2023 |
| record_format | arxiv |
| spellingShingle | The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models Chen, Xinyi Fernández, Raquel Pezzelle, Sandro Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what extent they handle basic linguistic constructions -- active-passive voice, coordination, and relative clauses -- that even preschool children can typically master. We present BLA, a novel, automatically constructed benchmark to evaluate multimodal models on these Basic Language Abilities. We show that different types of Transformer-based systems, such as CLIP, ViLBERT, and BLIP2, generally struggle with BLA in a zero-shot setting, in line with previous findings. Our experiments, in particular, show that most of the tested models only marginally benefit when fine-tuned or prompted with construction-specific samples. Yet, the generative BLIP2 shows promising trends, especially in an in-context learning setting. This opens the door to using BLA not only as an evaluation benchmark but also to improve models' basic language abilities. |
| title | The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models |
| topic | Computation and Language Artificial Intelligence Computer Vision and Pattern Recognition |
| url | https://arxiv.org/abs/2310.15061 |