Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Berger, Uri, Abend, Omri, Frermann, Lea, Stanovsky, Gabriel
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866912180374339584
author Berger, Uri
Abend, Omri
Frermann, Lea
Stanovsky, Gabriel
author_facet Berger, Uri
Abend, Omri
Frermann, Lea
Stanovsky, Gabriel
contents Incorporating automatically predicted human feedback into the process of training generative models has attracted substantial recent interest, while feedback at inference time has received less attention. The typical feedback at training time, i.e., preferences of choice given two samples, does not naturally transfer to the inference phase. We introduce a novel type of feedback -- caption reformulations -- and train models to mimic reformulation feedback based on human annotations. Our method does not require training the image captioning model itself, thereby demanding substantially less computational effort. We experiment with two types of reformulation feedback: first, we collect a dataset of human reformulations that correct errors in the generated captions. We find that incorporating reformulation models trained on this data into the inference phase of existing image captioning models results in improved captions, especially when the original captions are of low quality. We apply our method to non-English image captioning, a domain where robust models are less prevalent, and gain substantial improvement. Second, we apply reformulations to style transfer. Quantitative evaluations reveal state-of-the-art performance on German image captioning and English style transfer, while human validation with a detailed comparative framework exposes the specific axes of improvement.
format Preprint
id arxiv_https___arxiv_org_abs_2501_04513
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
Berger, Uri
Abend, Omri
Frermann, Lea
Stanovsky, Gabriel
Computer Vision and Pattern Recognition
Computation and Language
Incorporating automatically predicted human feedback into the process of training generative models has attracted substantial recent interest, while feedback at inference time has received less attention. The typical feedback at training time, i.e., preferences of choice given two samples, does not naturally transfer to the inference phase. We introduce a novel type of feedback -- caption reformulations -- and train models to mimic reformulation feedback based on human annotations. Our method does not require training the image captioning model itself, thereby demanding substantially less computational effort. We experiment with two types of reformulation feedback: first, we collect a dataset of human reformulations that correct errors in the generated captions. We find that incorporating reformulation models trained on this data into the inference phase of existing image captioning models results in improved captions, especially when the original captions are of low quality. We apply our method to non-English image captioning, a domain where robust models are less prevalent, and gain substantial improvement. Second, we apply reformulations to style transfer. Quantitative evaluations reveal state-of-the-art performance on German image captioning and English style transfer, while human validation with a detailed comparative framework exposes the specific axes of improvement.
title Improving Image Captioning by Mimicking Human Reformulation Feedback at Inference-time
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2501.04513