Grounding Language Models for Visual Entity Recognition

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Xiao, Zilin, Gong, Ming, Cascante-Bonilla, Paola, Zhang, Xingyao, Wu, Jie, Ordonez, Vicente
Format: Preprint
Published: 2024
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916336542679040
author Xiao, Zilin
Gong, Ming
Cascante-Bonilla, Paola
Zhang, Xingyao
Wu, Jie
Ordonez, Vicente
author_facet Xiao, Zilin
Gong, Ming
Cascante-Bonilla, Paola
Zhang, Xingyao
Wu, Jie
Ordonez, Vicente
contents We introduce AutoVER, an Autoregressive model for Visual Entity Recognition. Our model extends an autoregressive Multi-modal Large Language Model by employing retrieval augmented constrained generation. It mitigates low performance on out-of-domain entities while excelling in queries that require visually-situated reasoning. Our method learns to distinguish similar entities within a vast label space by contrastively training on hard negative pairs in parallel with a sequence-to-sequence objective without an external retriever. During inference, a list of retrieved candidate answers explicitly guides language generation by removing invalid decoding paths. The proposed method achieves significant improvements across different dataset splits in the recently proposed Oven-Wiki benchmark. Accuracy on the Entity seen split rises from 32.7% to 61.5%. It also demonstrates superior performance on the unseen and query splits by a substantial double-digit margin.
format Preprint
id arxiv_https___arxiv_org_abs_2402_18695
institution arXiv
publishDate 2024
record_format arxiv
spellingShingle Grounding Language Models for Visual Entity Recognition
Xiao, Zilin
Gong, Ming
Cascante-Bonilla, Paola
Zhang, Xingyao
Wu, Jie
Ordonez, Vicente
Computer Vision and Pattern Recognition
Computation and Language
We introduce AutoVER, an Autoregressive model for Visual Entity Recognition. Our model extends an autoregressive Multi-modal Large Language Model by employing retrieval augmented constrained generation. It mitigates low performance on out-of-domain entities while excelling in queries that require visually-situated reasoning. Our method learns to distinguish similar entities within a vast label space by contrastively training on hard negative pairs in parallel with a sequence-to-sequence objective without an external retriever. During inference, a list of retrieved candidate answers explicitly guides language generation by removing invalid decoding paths. The proposed method achieves significant improvements across different dataset splits in the recently proposed Oven-Wiki benchmark. Accuracy on the Entity seen split rises from 32.7% to 61.5%. It also demonstrates superior performance on the unseen and query splits by a substantial double-digit margin.
title Grounding Language Models for Visual Entity Recognition
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2402.18695