Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction

Fuente: arXiv
Salvato in:
Dettagli Bibliografici
Autori principali: Fonseca, Rui, Martins, Bruno, Rocha, Gil
Natura: Preprint
Pubblicazione: 2025
Soggetti:
Accesso online:
Tags: Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
_version_ 1866909943270998016
author Fonseca, Rui
Martins, Bruno
Rocha, Gil
author_facet Fonseca, Rui
Martins, Bruno
Rocha, Gil
contents Image captioning has drawn considerable attention from the natural language processing and computer vision fields. Aiming to reduce the reliance on curated data, several studies have explored image captioning without any humanly-annotated image-text pairs for training, although existing methods are still outperformed by fully supervised approaches. This paper proposes TOMCap, i.e., an improved text-only training method that performs captioning without the need for aligned image-caption pairs. The method is based on prompting a pre-trained language model decoder with information derived from a CLIP representation, after undergoing a process to reduce the modality gap. We specifically tested the combined use of retrieved examples of captions, and latent vector representations, to guide the generation process. Through extensive experiments, we show that TOMCap outperforms other training-free and text-only methods. We also analyze the impact of different choices regarding the configuration of the retrieval-augmentation and modality gap reduction components.
format Preprint
id arxiv_https___arxiv_org_abs_2512_04309
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
Fonseca, Rui
Martins, Bruno
Rocha, Gil
Computer Vision and Pattern Recognition
Computation and Language
Image captioning has drawn considerable attention from the natural language processing and computer vision fields. Aiming to reduce the reliance on curated data, several studies have explored image captioning without any humanly-annotated image-text pairs for training, although existing methods are still outperformed by fully supervised approaches. This paper proposes TOMCap, i.e., an improved text-only training method that performs captioning without the need for aligned image-caption pairs. The method is based on prompting a pre-trained language model decoder with information derived from a CLIP representation, after undergoing a process to reduce the modality gap. We specifically tested the combined use of retrieved examples of captions, and latent vector representations, to guide the generation process. Through extensive experiments, we show that TOMCap outperforms other training-free and text-only methods. We also analyze the impact of different choices regarding the configuration of the retrieval-augmentation and modality gap reduction components.
title Text-Only Training for Image Captioning with Retrieval Augmentation and Modality Gap Correction
topic Computer Vision and Pattern Recognition
Computation and Language
url https://arxiv.org/abs/2512.04309