FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ruan, Jiacheng, Yang, Yebin, Lin, Zehao, Feng, Yuchen, Xiong, Feiyu, Tang, Zeyun, Li, Zhiyu
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866917847003824128
author Ruan, Jiacheng
Yang, Yebin
Lin, Zehao
Feng, Yuchen
Xiong, Feiyu
Tang, Zeyun
Li, Zhiyu
author_facet Ruan, Jiacheng
Yang, Yebin
Lin, Zehao
Feng, Yuchen
Xiong, Feiyu
Tang, Zeyun
Li, Zhiyu
contents Benefiting from the revolutionary advances in large language models (LLMs) and foundational vision models, large vision-language models (LVLMs) have also made significant progress. However, current benchmarks focus on tasks that evaluating only a single aspect of LVLM capabilities (e.g., recognition, detection, understanding). These tasks fail to fully demonstrate LVLMs' potential in complex application scenarios. To comprehensively assess the performance of existing LVLMs, we propose a more challenging task called the Flow Text with Image Insertion task (FTII). This task requires LVLMs to simultaneously possess outstanding abilities in image comprehension, instruction understanding, and long-text interpretation. Specifically, given several text paragraphs and a set of candidate images, as the text paragraphs accumulate, the LVLMs are required to select the most suitable image from the candidates to insert after the corresponding paragraph. Constructing a benchmark for such a task is highly challenging, particularly in determining the sequence of flowing text and images. To address this challenge, we turn to professional news reports, which naturally contain a gold standard for image-text sequences. Based on this, we introduce the Flow Text with Image Insertion Benchmark (FTII-Bench), which includes 318 high-quality Chinese image-text news articles and 307 high-quality English image-text news articles, covering 10 different news domains. Using these 625 high-quality articles, we construct problems of two different types with multiple levels of difficulty. Furthermore, we establish two different evaluation pipelines based on the CLIP model and existing LVLMs. We evaluate 9 open-source and 2 closed-source LVLMs as well as 2 CLIP-based models. Results indicate that even the most advanced models (e.g., GPT-4o) face significant challenges when tackling the FTII task.
format Preprint
id arxiv_https___arxiv_org_abs_2410_12564
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion
Ruan, Jiacheng
Yang, Yebin
Lin, Zehao
Feng, Yuchen
Xiong, Feiyu
Tang, Zeyun
Li, Zhiyu
Computer Vision and Pattern Recognition
Benefiting from the revolutionary advances in large language models (LLMs) and foundational vision models, large vision-language models (LVLMs) have also made significant progress. However, current benchmarks focus on tasks that evaluating only a single aspect of LVLM capabilities (e.g., recognition, detection, understanding). These tasks fail to fully demonstrate LVLMs' potential in complex application scenarios. To comprehensively assess the performance of existing LVLMs, we propose a more challenging task called the Flow Text with Image Insertion task (FTII). This task requires LVLMs to simultaneously possess outstanding abilities in image comprehension, instruction understanding, and long-text interpretation. Specifically, given several text paragraphs and a set of candidate images, as the text paragraphs accumulate, the LVLMs are required to select the most suitable image from the candidates to insert after the corresponding paragraph. Constructing a benchmark for such a task is highly challenging, particularly in determining the sequence of flowing text and images. To address this challenge, we turn to professional news reports, which naturally contain a gold standard for image-text sequences. Based on this, we introduce the Flow Text with Image Insertion Benchmark (FTII-Bench), which includes 318 high-quality Chinese image-text news articles and 307 high-quality English image-text news articles, covering 10 different news domains. Using these 625 high-quality articles, we construct problems of two different types with multiple levels of difficulty. Furthermore, we establish two different evaluation pipelines based on the CLIP model and existing LVLMs. We evaluate 9 open-source and 2 closed-source LVLMs as well as 2 CLIP-based models. Results indicate that even the most advanced models (e.g., GPT-4o) face significant challenges when tackling the FTII task.
title FTII-Bench: A Comprehensive Multimodal Benchmark for Flow Text with Image Insertion
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.12564