SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Truong, Khang, Pham, Lam, Tang, Hieu, Lampert, Jasmin, Boyer, Martin, Phan, Son, Nguyen, Truong
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918095783723008
author Truong, Khang
Pham, Lam
Tang, Hieu
Lampert, Jasmin
Boyer, Martin
Phan, Son
Nguyen, Truong
author_facet Truong, Khang
Pham, Lam
Tang, Hieu
Lampert, Jasmin
Boyer, Martin
Phan, Son
Nguyen, Truong
contents Image captioning has emerged as a crucial task in the intersection of computer vision and natural language processing, enabling automated generation of descriptive text from visual content. In the context of remote sensing, image captioning plays a significant role in interpreting vast and complex satellite imagery, aiding applications such as environmental monitoring, disaster assessment, and urban planning. This motivates us, in this paper, to present a transformer based network architecture for remote sensing image captioning (RSIC) in which multiple techniques of Static Expansion, Memory-Augmented Self-Attention, Mesh Transformer are evaluated and integrated. We evaluate our proposed models using two benchmark remote sensing image datasets of UCM-Caption and NWPU-Caption. Our best model outperforms the state-of-the-art systems on most of evaluation metrics, which demonstrates potential to apply for real-life remote sensing image systems.
format Preprint
id arxiv_https___arxiv_org_abs_2507_12845
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
Truong, Khang
Pham, Lam
Tang, Hieu
Lampert, Jasmin
Boyer, Martin
Phan, Son
Nguyen, Truong
Computer Vision and Pattern Recognition
Artificial Intelligence
Image captioning has emerged as a crucial task in the intersection of computer vision and natural language processing, enabling automated generation of descriptive text from visual content. In the context of remote sensing, image captioning plays a significant role in interpreting vast and complex satellite imagery, aiding applications such as environmental monitoring, disaster assessment, and urban planning. This motivates us, in this paper, to present a transformer based network architecture for remote sensing image captioning (RSIC) in which multiple techniques of Static Expansion, Memory-Augmented Self-Attention, Mesh Transformer are evaluated and integrated. We evaluate our proposed models using two benchmark remote sensing image datasets of UCM-Caption and NWPU-Caption. Our best model outperforms the state-of-the-art systems on most of evaluation metrics, which demonstrates potential to apply for real-life remote sensing image systems.
title SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2507.12845