Mitigating Open-Vocabulary Caption Hallucinations

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Ben-Kish, Assaf, Yanuka, Moran, Alper, Morris, Giryes, Raja, Averbuch-Elor, Hadar
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866916580343939072
author Ben-Kish, Assaf
Yanuka, Moran
Alper, Morris
Giryes, Raja
Averbuch-Elor, Hadar
author_facet Ben-Kish, Assaf
Yanuka, Moran
Alper, Morris
Giryes, Raja
Averbuch-Elor, Hadar
contents While recent years have seen rapid progress in image-conditioned text generation, image captioning still suffers from the fundamental issue of hallucinations, namely, the generation of spurious details that cannot be inferred from the given image. Existing methods largely use closed-vocabulary object lists to mitigate or evaluate hallucinations in image captioning, ignoring the long-tailed nature of hallucinations that occur in practice. To this end, we propose a framework for addressing hallucinations in image captioning in the open-vocabulary setting. Our framework includes a new benchmark, OpenCHAIR, that leverages generative foundation models to evaluate open-vocabulary object hallucinations for image captioning, surpassing the popular and similarly-sized CHAIR benchmark in both diversity and accuracy. Furthermore, to mitigate open-vocabulary hallucinations without using a closed object list, we propose MOCHa, an approach harnessing advancements in reinforcement learning. Our multi-objective reward function explicitly targets the trade-off between fidelity and adequacy in generations without requiring any strong supervision. MOCHa improves a large variety of image captioning models, as captured by our OpenCHAIR benchmark and other existing metrics. Code and models can be found at: https://github.com/assafbk/mocha_code
format Preprint
id arxiv_https___arxiv_org_abs_2312_03631
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Mitigating Open-Vocabulary Caption Hallucinations
Ben-Kish, Assaf
Yanuka, Moran
Alper, Morris
Giryes, Raja
Averbuch-Elor, Hadar
Computer Vision and Pattern Recognition
Artificial Intelligence
While recent years have seen rapid progress in image-conditioned text generation, image captioning still suffers from the fundamental issue of hallucinations, namely, the generation of spurious details that cannot be inferred from the given image. Existing methods largely use closed-vocabulary object lists to mitigate or evaluate hallucinations in image captioning, ignoring the long-tailed nature of hallucinations that occur in practice. To this end, we propose a framework for addressing hallucinations in image captioning in the open-vocabulary setting. Our framework includes a new benchmark, OpenCHAIR, that leverages generative foundation models to evaluate open-vocabulary object hallucinations for image captioning, surpassing the popular and similarly-sized CHAIR benchmark in both diversity and accuracy. Furthermore, to mitigate open-vocabulary hallucinations without using a closed object list, we propose MOCHa, an approach harnessing advancements in reinforcement learning. Our multi-objective reward function explicitly targets the trade-off between fidelity and adequacy in generations without requiring any strong supervision. MOCHa improves a large variety of image captioning models, as captured by our OpenCHAIR benchmark and other existing metrics. Code and models can be found at: https://github.com/assafbk/mocha_code
title Mitigating Open-Vocabulary Caption Hallucinations
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2312.03631