EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Li, Jiaxuan, Vo, Duc Minh, Sugimoto, Akihiro, Nakayama, Hideki
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914742963011584
author Li, Jiaxuan
Vo, Duc Minh
Sugimoto, Akihiro
Nakayama, Hideki
author_facet Li, Jiaxuan
Vo, Duc Minh
Sugimoto, Akihiro
Nakayama, Hideki
contents Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large amounts of data and/or scaling up network parameters, we introduce a highly effective retrieval-augmented image captioning method that prompts LLMs with object names retrieved from External Visual--name memory (EVCap). We build ever-changing object knowledge memory using objects' visuals and names, enabling us to (i) update the memory at a minimal cost and (ii) effortlessly augment LLMs with retrieved object names by utilizing a lightweight and fast-to-train model. Our model, which was trained only on the COCO dataset, can adapt to out-of-domain without requiring additional fine-tuning or re-training. Our experiments conducted on benchmarks and synthetic commonsense-violating data show that EVCap, with only 3.97M trainable parameters, exhibits superior performance compared to other methods based on frozen pre-trained LLMs. Its performance is also competitive to specialist SOTAs that require extensive training.
format Preprint
id arxiv_https___arxiv_org_abs_2311_15879
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
Li, Jiaxuan
Vo, Duc Minh
Sugimoto, Akihiro
Nakayama, Hideki
Computer Vision and Pattern Recognition
Large language models (LLMs)-based image captioning has the capability of describing objects not explicitly observed in training data; yet novel objects occur frequently, necessitating the requirement of sustaining up-to-date object knowledge for open-world comprehension. Instead of relying on large amounts of data and/or scaling up network parameters, we introduce a highly effective retrieval-augmented image captioning method that prompts LLMs with object names retrieved from External Visual--name memory (EVCap). We build ever-changing object knowledge memory using objects' visuals and names, enabling us to (i) update the memory at a minimal cost and (ii) effortlessly augment LLMs with retrieved object names by utilizing a lightweight and fast-to-train model. Our model, which was trained only on the COCO dataset, can adapt to out-of-domain without requiring additional fine-tuning or re-training. Our experiments conducted on benchmarks and synthetic commonsense-violating data show that EVCap, with only 3.97M trainable parameters, exhibits superior performance compared to other methods based on frozen pre-trained LLMs. Its performance is also competitive to specialist SOTAs that require extensive training.
title EVCap: Retrieval-Augmented Image Captioning with External Visual-Name Memory for Open-World Comprehension
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2311.15879