Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Peng, Wujian, Xie, Sicheng, You, Zuyao, Lan, Shiyi, Wu, Zuxuan
Formato: Preprint
Publicado: 2023
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866913291636310016
author Peng, Wujian
Xie, Sicheng
You, Zuyao
Lan, Shiyi
Wu, Zuxuan
author_facet Peng, Wujian
Xie, Sicheng
You, Zuyao
Lan, Shiyi
Wu, Zuxuan
contents Vision language models (VLM) have demonstrated remarkable performance across various downstream tasks. However, understanding fine-grained visual-linguistic concepts, such as attributes and inter-object relationships, remains a significant challenge. While several benchmarks aim to evaluate VLMs in finer granularity, their primary focus remains on the linguistic aspect, neglecting the visual dimension. Here, we highlight the importance of evaluating VLMs from both a textual and visual perspective. We introduce a progressive pipeline to synthesize images that vary in a specific attribute while ensuring consistency in all other aspects. Utilizing this data engine, we carefully design a benchmark, SPEC, to diagnose the comprehension of object size, position, existence, and count. Subsequently, we conduct a thorough evaluation of four leading VLMs on SPEC. Surprisingly, their performance is close to random guess, revealing significant limitations. With this in mind, we propose a simple yet effective approach to optimize VLMs in fine-grained understanding, achieving significant improvements on SPEC without compromising the zero-shot performance. Results on two additional fine-grained benchmarks also show consistent improvements, further validating the transferability of our approach. Code and data are available at https://github.com/wjpoom/SPEC.
format Preprint
id arxiv_https___arxiv_org_abs_2312_00081
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding
Peng, Wujian
Xie, Sicheng
You, Zuyao
Lan, Shiyi
Wu, Zuxuan
Computer Vision and Pattern Recognition
Vision language models (VLM) have demonstrated remarkable performance across various downstream tasks. However, understanding fine-grained visual-linguistic concepts, such as attributes and inter-object relationships, remains a significant challenge. While several benchmarks aim to evaluate VLMs in finer granularity, their primary focus remains on the linguistic aspect, neglecting the visual dimension. Here, we highlight the importance of evaluating VLMs from both a textual and visual perspective. We introduce a progressive pipeline to synthesize images that vary in a specific attribute while ensuring consistency in all other aspects. Utilizing this data engine, we carefully design a benchmark, SPEC, to diagnose the comprehension of object size, position, existence, and count. Subsequently, we conduct a thorough evaluation of four leading VLMs on SPEC. Surprisingly, their performance is close to random guess, revealing significant limitations. With this in mind, we propose a simple yet effective approach to optimize VLMs in fine-grained understanding, achieving significant improvements on SPEC without compromising the zero-shot performance. Results on two additional fine-grained benchmarks also show consistent improvements, further validating the transferability of our approach. Code and data are available at https://github.com/wjpoom/SPEC.
title Synthesize, Diagnose, and Optimize: Towards Fine-Grained Vision-Language Understanding
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2312.00081