Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Albadarneh, Israa A., Hammo, Bassam H., Al-Kadi, Omar S.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915328751042560
author Albadarneh, Israa A.
Hammo, Bassam H.
Al-Kadi, Omar S.
author_facet Albadarneh, Israa A.
Hammo, Bassam H.
Al-Kadi, Omar S.
contents Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly improved caption generation by leveraging attention mechanisms for better scene understanding. While various surveys have explored deep learning-based approaches for image captioning, few have comprehensively analyzed attention-based transformer models across multiple languages. This survey reviews attention-based image captioning models, categorizing them into transformer-based, deep learning-based, and hybrid approaches. It explores benchmark datasets, discusses evaluation metrics such as BLEU, METEOR, CIDEr, and ROUGE, and highlights challenges in multilingual captioning. Additionally, this paper identifies key limitations in current models, including semantic inconsistencies, data scarcity in non-English languages, and limitations in reasoning ability. Finally, we outline future research directions, such as multimodal learning, real-time applications in AI-powered assistants, healthcare, and forensic analysis. This survey serves as a comprehensive reference for researchers aiming to advance the field of attention-based image captioning.
format Preprint
id arxiv_https___arxiv_org_abs_2506_05399
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
Albadarneh, Israa A.
Hammo, Bassam H.
Al-Kadi, Omar S.
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Image captioning involves generating textual descriptions from input images, bridging the gap between computer vision and natural language processing. Recent advancements in transformer-based models have significantly improved caption generation by leveraging attention mechanisms for better scene understanding. While various surveys have explored deep learning-based approaches for image captioning, few have comprehensively analyzed attention-based transformer models across multiple languages. This survey reviews attention-based image captioning models, categorizing them into transformer-based, deep learning-based, and hybrid approaches. It explores benchmark datasets, discusses evaluation metrics such as BLEU, METEOR, CIDEr, and ROUGE, and highlights challenges in multilingual captioning. Additionally, this paper identifies key limitations in current models, including semantic inconsistencies, data scarcity in non-English languages, and limitations in reasoning ability. Finally, we outline future research directions, such as multimodal learning, real-time applications in AI-powered assistants, healthcare, and forensic analysis. This survey serves as a comprehensive reference for researchers aiming to advance the field of attention-based image captioning.
title Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2506.05399