GenCeption: Evaluate Vision LLMs with Unlabeled Unimodal Data

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Cao, Lele, Buchner, Valentin, Senane, Zineb, Yang, Fangkai
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916644747476992
author Cao, Lele
Buchner, Valentin
Senane, Zineb
Yang, Fangkai
author_facet Cao, Lele
Buchner, Valentin
Senane, Zineb
Yang, Fangkai
contents Multimodal Large Language Models (MLLMs) are typically assessed using expensive annotated multimodal benchmarks, which often lag behind the rapidly evolving demands of MLLM evaluation. This paper outlines and validates GenCeption, a novel, annotation-free evaluation method that requires only unimodal data to measure inter-modality semantic coherence and inversely assesses MLLMs' tendency to hallucinate. This approach eliminates the need for costly data annotation, minimizes the risk of training data contamination, is expected to result in slower benchmark saturation, and avoids the illusion of emerging abilities. Inspired by the DrawCeption game, GenCeption begins with a non-textual sample and proceeds through iterative description and generation steps. The semantic drift across iterations is quantified using the GC@T metric. While GenCeption is principally applicable to MLLMs across various modalities, this paper focuses on its implementation and validation for Vision LLMs (VLLMs). Based on the GenCeption method, we establish the MMECeption benchmark for evaluating VLLMs, and compare the performance of several popular VLLMs and human annotators. Our empirical results validate GenCeption's effectiveness, demonstrating strong correlations with established VLLM benchmarks. VLLMs still significantly lag behind human performance and struggle especially with text-intensive tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2402_14973
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle GenCeption: Evaluate Vision LLMs with Unlabeled Unimodal Data
Cao, Lele
Buchner, Valentin
Senane, Zineb
Yang, Fangkai
Computation and Language
Artificial Intelligence
Machine Learning
I.7; I.4
Multimodal Large Language Models (MLLMs) are typically assessed using expensive annotated multimodal benchmarks, which often lag behind the rapidly evolving demands of MLLM evaluation. This paper outlines and validates GenCeption, a novel, annotation-free evaluation method that requires only unimodal data to measure inter-modality semantic coherence and inversely assesses MLLMs' tendency to hallucinate. This approach eliminates the need for costly data annotation, minimizes the risk of training data contamination, is expected to result in slower benchmark saturation, and avoids the illusion of emerging abilities. Inspired by the DrawCeption game, GenCeption begins with a non-textual sample and proceeds through iterative description and generation steps. The semantic drift across iterations is quantified using the GC@T metric. While GenCeption is principally applicable to MLLMs across various modalities, this paper focuses on its implementation and validation for Vision LLMs (VLLMs). Based on the GenCeption method, we establish the MMECeption benchmark for evaluating VLLMs, and compare the performance of several popular VLLMs and human annotators. Our empirical results validate GenCeption's effectiveness, demonstrating strong correlations with established VLLM benchmarks. VLLMs still significantly lag behind human performance and struggle especially with text-intensive tasks.
title GenCeption: Evaluate Vision LLMs with Unlabeled Unimodal Data
topic Computation and Language
Artificial Intelligence
Machine Learning
I.7; I.4
url https://arxiv.org/abs/2402.14973