Visually-Aware Context Modeling for News Image Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Qu, Tingyu, Tuytelaars, Tinne, Moens, Marie-Francine
Format: Preprint
Published: 2023
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866911806140710912
author Qu, Tingyu
Tuytelaars, Tinne
Moens, Marie-Francine
author_facet Qu, Tingyu
Tuytelaars, Tinne
Moens, Marie-Francine
contents News Image Captioning aims to create captions from news articles and images, emphasizing the connection between textual context and visual elements. Recognizing the significance of human faces in news images and the face-name co-occurrence pattern in existing datasets, we propose a face-naming module for learning better name embeddings. Apart from names, which can be directly linked to an image area (faces), news image captions mostly contain context information that can only be found in the article. We design a retrieval strategy using CLIP to retrieve sentences that are semantically close to the image, mimicking human thought process of linking articles to images. Furthermore, to tackle the problem of the imbalanced proportion of article context and image context in captions, we introduce a simple yet effective method Contrasting with Language Model backbone (CoLaM) to the training pipeline. We conduct extensive experiments to demonstrate the efficacy of our framework. We out-perform the previous state-of-the-art (without external data) by 7.97/5.80 CIDEr scores on GoodNews/NYTimes800k. Our code is available at https://github.com/tingyu215/VACNIC.
format Preprint
id arxiv_https___arxiv_org_abs_2308_08325
institution arXiv
publishDate 2023
record_format arxiv
spellingShingle Visually-Aware Context Modeling for News Image Captioning
Qu, Tingyu
Tuytelaars, Tinne
Moens, Marie-Francine
Computer Vision and Pattern Recognition
News Image Captioning aims to create captions from news articles and images, emphasizing the connection between textual context and visual elements. Recognizing the significance of human faces in news images and the face-name co-occurrence pattern in existing datasets, we propose a face-naming module for learning better name embeddings. Apart from names, which can be directly linked to an image area (faces), news image captions mostly contain context information that can only be found in the article. We design a retrieval strategy using CLIP to retrieve sentences that are semantically close to the image, mimicking human thought process of linking articles to images. Furthermore, to tackle the problem of the imbalanced proportion of article context and image context in captions, we introduce a simple yet effective method Contrasting with Language Model backbone (CoLaM) to the training pipeline. We conduct extensive experiments to demonstrate the efficacy of our framework. We out-perform the previous state-of-the-art (without external data) by 7.97/5.80 CIDEr scores on GoodNews/NYTimes800k. Our code is available at https://github.com/tingyu215/VACNIC.
title Visually-Aware Context Modeling for News Image Captioning
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2308.08325