Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Caron, Mathilde, Fathi, Alireza, Schmid, Cordelia, Iscen, Ahmet
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913568550551552
author Caron, Mathilde
Fathi, Alireza
Schmid, Cordelia
Iscen, Ahmet
author_facet Caron, Mathilde
Fathi, Alireza
Schmid, Cordelia
Iscen, Ahmet
contents Web-scale visual entity recognition, the task of associating images with their corresponding entities within vast knowledge bases like Wikipedia, presents significant challenges due to the lack of clean, large-scale training data. In this paper, we propose a novel methodology to curate such a dataset, leveraging a multimodal large language model (LLM) for label verification, metadata generation, and rationale explanation. Instead of relying on the multimodal LLM to directly annotate data, which we found to be suboptimal, we prompt it to reason about potential candidate entity labels by accessing additional contextually relevant information (such as Wikipedia), resulting in more accurate annotations. We further use the multimodal LLM to enrich the dataset by generating question-answer pairs and a grounded finegrained textual description (referred to as "rationale") that explains the connection between images and their assigned entities. Experiments demonstrate that models trained on this automatically curated data achieve state-of-the-art performance on web-scale visual entity recognition tasks (e.g. +6.9% improvement in OVEN entity task), underscoring the importance of high-quality training data in this domain.
format Preprint
id arxiv_https___arxiv_org_abs_2410_23676
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
Caron, Mathilde
Fathi, Alireza
Schmid, Cordelia
Iscen, Ahmet
Computer Vision and Pattern Recognition
Web-scale visual entity recognition, the task of associating images with their corresponding entities within vast knowledge bases like Wikipedia, presents significant challenges due to the lack of clean, large-scale training data. In this paper, we propose a novel methodology to curate such a dataset, leveraging a multimodal large language model (LLM) for label verification, metadata generation, and rationale explanation. Instead of relying on the multimodal LLM to directly annotate data, which we found to be suboptimal, we prompt it to reason about potential candidate entity labels by accessing additional contextually relevant information (such as Wikipedia), resulting in more accurate annotations. We further use the multimodal LLM to enrich the dataset by generating question-answer pairs and a grounded finegrained textual description (referred to as "rationale") that explains the connection between images and their assigned entities. Experiments demonstrate that models trained on this automatically curated data achieve state-of-the-art performance on web-scale visual entity recognition tasks (e.g. +6.9% improvement in OVEN entity task), underscoring the importance of high-quality training data in this domain.
title Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2410.23676