AGIC: Attention-Guided Image Captioning to Improve Caption Relevance

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Teja, L. D. M. S. Sai, Urlana, Ashok, Mishra, Pruthwik
Format: Preprint
Veröffentlicht: 2025
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866913982700322816
author Teja, L. D. M. S. Sai
Urlana, Ashok
Mishra, Pruthwik
author_facet Teja, L. D. M. S. Sai
Urlana, Ashok
Mishra, Pruthwik
contents Despite significant progress in image captioning, generating accurate and descriptive captions remains a long-standing challenge. In this study, we propose Attention-Guided Image Captioning (AGIC), which amplifies salient visual regions directly in the feature space to guide caption generation. We further introduce a hybrid decoding strategy that combines deterministic and probabilistic sampling to balance fluency and diversity. To evaluate AGIC, we conduct extensive experiments on the Flickr8k and Flickr30k datasets. The results show that AGIC matches or surpasses several state-of-the-art models while achieving faster inference. Moreover, AGIC demonstrates strong performance across multiple evaluation metrics, offering a scalable and interpretable solution for image captioning.
format Preprint
id arxiv_https___arxiv_org_abs_2508_06853
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
Teja, L. D. M. S. Sai
Urlana, Ashok
Mishra, Pruthwik
Computer Vision and Pattern Recognition
Artificial Intelligence
Despite significant progress in image captioning, generating accurate and descriptive captions remains a long-standing challenge. In this study, we propose Attention-Guided Image Captioning (AGIC), which amplifies salient visual regions directly in the feature space to guide caption generation. We further introduce a hybrid decoding strategy that combines deterministic and probabilistic sampling to balance fluency and diversity. To evaluate AGIC, we conduct extensive experiments on the Flickr8k and Flickr30k datasets. The results show that AGIC matches or surpasses several state-of-the-art models while achieving faster inference. Moreover, AGIC demonstrates strong performance across multiple evaluation metrics, offering a scalable and interpretable solution for image captioning.
title AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2508.06853