UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Pi, Xinyu, Wu, Mingyuan, Jiang, Jize, Zheng, Haozhen, Tian, Beitong, Zhai, Chengxiang, Nahrstedt, Klara, Hu, Zhiting
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911968567230464
author Pi, Xinyu
Wu, Mingyuan
Jiang, Jize
Zheng, Haozhen
Tian, Beitong
Zhai, Chengxiang
Nahrstedt, Klara
Hu, Zhiting
author_facet Pi, Xinyu
Wu, Mingyuan
Jiang, Jize
Zheng, Haozhen
Tian, Beitong
Zhai, Chengxiang
Nahrstedt, Klara
Hu, Zhiting
contents Smaller-scale Vision-Langauge Models (VLMs) often claim to perform on par with larger models in general-domain visual grounding and question-answering benchmarks while offering advantages in computational efficiency and storage. However, their ability to handle rare objects, which fall into the long tail of data distributions, is less understood. To rigorously evaluate this aspect, we introduce the "Uncontextualized Uncommon Objects" (UOUO) benchmark. This benchmark focuses on systematically testing VLMs with both large and small parameter counts on rare and specialized objects. Our comprehensive analysis reveals that while smaller VLMs maintain competitive performance on common datasets, they significantly underperform on tasks involving uncommon objects. We also propose an advanced, scalable pipeline for data collection and cleaning, ensuring the UOUO benchmark provides high-quality, challenging instances. These findings highlight the need to consider long-tail distributions when assessing the true capabilities of VLMs.
format Preprint
id arxiv_https___arxiv_org_abs_2407_18391
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
Pi, Xinyu
Wu, Mingyuan
Jiang, Jize
Zheng, Haozhen
Tian, Beitong
Zhai, Chengxiang
Nahrstedt, Klara
Hu, Zhiting
Computer Vision and Pattern Recognition
Smaller-scale Vision-Langauge Models (VLMs) often claim to perform on par with larger models in general-domain visual grounding and question-answering benchmarks while offering advantages in computational efficiency and storage. However, their ability to handle rare objects, which fall into the long tail of data distributions, is less understood. To rigorously evaluate this aspect, we introduce the "Uncontextualized Uncommon Objects" (UOUO) benchmark. This benchmark focuses on systematically testing VLMs with both large and small parameter counts on rare and specialized objects. Our comprehensive analysis reveals that while smaller VLMs maintain competitive performance on common datasets, they significantly underperform on tasks involving uncommon objects. We also propose an advanced, scalable pipeline for data collection and cleaning, ensuring the UOUO benchmark provides high-quality, challenging instances. These findings highlight the need to consider long-tail distributions when assessing the true capabilities of VLMs.
title UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2407.18391