VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Parcalabescu, Letitia, Cafagna, Michele, Muradjan, Lilitta, Frank, Anette, Calixto, Iacer, Gatt, Albert
Formato: Preprint
Publicado: 2021
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866916119821942784
author Parcalabescu, Letitia
Cafagna, Michele
Muradjan, Lilitta
Frank, Anette
Calixto, Iacer
Gatt, Albert
author_facet Parcalabescu, Letitia
Cafagna, Michele
Muradjan, Lilitta
Frank, Anette
Calixto, Iacer
Gatt, Albert
contents We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena. VALSE offers a suite of six tests covering various linguistic constructs. Solving these requires models to ground linguistic phenomena in the visual modality, allowing more fine-grained evaluations than hitherto possible. We build VALSE using methods that support the construction of valid foils, and report results from evaluating five widely-used V&L models. Our experiments suggest that current models have considerable difficulty addressing most phenomena. Hence, we expect VALSE to serve as an important benchmark to measure future progress of pretrained V&L models from a linguistic perspective, complementing the canonical task-centred V&L evaluations.
format Preprint
id arxiv_https___arxiv_org_abs_2112_07566
institution arXiv
publishDate 2021
record_format arxiv
spellingShingle VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
Parcalabescu, Letitia
Cafagna, Michele
Muradjan, Lilitta
Frank, Anette
Calixto, Iacer
Gatt, Albert
Computation and Language
Computer Vision and Pattern Recognition
68Txx
I.2.7; I.2.10
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena. VALSE offers a suite of six tests covering various linguistic constructs. Solving these requires models to ground linguistic phenomena in the visual modality, allowing more fine-grained evaluations than hitherto possible. We build VALSE using methods that support the construction of valid foils, and report results from evaluating five widely-used V&L models. Our experiments suggest that current models have considerable difficulty addressing most phenomena. Hence, we expect VALSE to serve as an important benchmark to measure future progress of pretrained V&L models from a linguistic perspective, complementing the canonical task-centred V&L evaluations.
title VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
topic Computation and Language
Computer Vision and Pattern Recognition
68Txx
I.2.7; I.2.10
url https://arxiv.org/abs/2112.07566