Natural Language Inference Improves Compositionality in Vision-Language Models

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Cascante-Bonilla, Paola, Hou, Yu, Cao, Yang Trista, Daumé III, Hal, Rudinger, Rachel
Natura: Preprint
Pubblicazione: 2024
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866913567172722688
author Cascante-Bonilla, Paola
Hou, Yu
Cao, Yang Trista
Daumé III, Hal
Rudinger, Rachel
author_facet Cascante-Bonilla, Paola
Hou, Yu
Cao, Yang Trista
Daumé III, Hal
Rudinger, Rachel
contents Compositional reasoning in Vision-Language Models (VLMs) remains challenging as these models often struggle to relate objects, attributes, and spatial relationships. Recent methods aim to address these limitations by relying on the semantics of the textual description, using Large Language Models (LLMs) to break them down into subsets of questions and answers. However, these methods primarily operate on the surface level, failing to incorporate deeper lexical understanding while introducing incorrect assumptions generated by the LLM. In response to these issues, we present Caption Expansion with Contradictions and Entailments (CECE), a principled approach that leverages Natural Language Inference (NLI) to generate entailments and contradictions from a given premise. CECE produces lexically diverse sentences while maintaining their core meaning. Through extensive experiments, we show that CECE enhances interpretability and reduces overreliance on biased or superficial features. By balancing CECE along the original premise, we achieve significant improvements over previous methods without requiring additional fine-tuning, producing state-of-the-art results on benchmarks that score agreement with human judgments for image-text alignment, and achieving an increase in performance on Winoground of +19.2% (group score) and +12.9% on EqBen (group score) over the best prior work (finetuned with targeted data).
format Preprint
id arxiv_https___arxiv_org_abs_2410_22315
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Natural Language Inference Improves Compositionality in Vision-Language Models
Cascante-Bonilla, Paola
Hou, Yu
Cao, Yang Trista
Daumé III, Hal
Rudinger, Rachel
Computation and Language
Computer Vision and Pattern Recognition
Compositional reasoning in Vision-Language Models (VLMs) remains challenging as these models often struggle to relate objects, attributes, and spatial relationships. Recent methods aim to address these limitations by relying on the semantics of the textual description, using Large Language Models (LLMs) to break them down into subsets of questions and answers. However, these methods primarily operate on the surface level, failing to incorporate deeper lexical understanding while introducing incorrect assumptions generated by the LLM. In response to these issues, we present Caption Expansion with Contradictions and Entailments (CECE), a principled approach that leverages Natural Language Inference (NLI) to generate entailments and contradictions from a given premise. CECE produces lexically diverse sentences while maintaining their core meaning. Through extensive experiments, we show that CECE enhances interpretability and reduces overreliance on biased or superficial features. By balancing CECE along the original premise, we achieve significant improvements over previous methods without requiring additional fine-tuning, producing state-of-the-art results on benchmarks that score agreement with human judgments for image-text alignment, and achieving an increase in performance on Winoground of +19.2% (group score) and +12.9% on EqBen (group score) over the best prior work (finetuned with targeted data).
title Natural Language Inference Improves Compositionality in Vision-Language Models
topic Computation and Language
Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.22315