Saved in:
| Main Authors: | Zhang, Xu, Fang, Jiabin, Ding, Zhuoming, Yuan, Jin, Liu, Xuan, Zhang, Qianjun, Li, Zhiyong |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.11680 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
by: Zhang, Xu, et al.
Published: (2026)
by: Zhang, Xu, et al.
Published: (2026)
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
by: Zhang, Xu, et al.
Published: (2023)
by: Zhang, Xu, et al.
Published: (2023)
CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding
by: Zhang, Mingming, et al.
Published: (2023)
by: Zhang, Mingming, et al.
Published: (2023)
Learning a Cross-modality Anomaly Detector for Remote Sensing Imagery
by: Li, Jingtao, et al.
Published: (2023)
by: Li, Jingtao, et al.
Published: (2023)
Knowledge-aware Visual Question Generation for Remote Sensing Images
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
CADFormer: Fine-Grained Cross-modal Alignment and Decoding Transformer for Referring Remote Sensing Image Segmentation
by: Liu, Maofu, et al.
Published: (2025)
by: Liu, Maofu, et al.
Published: (2025)
Observe Less, Understand More: Cost-aware Cross-scale Observation for Remote Sensing Understanding
by: Xie, Zhenghao, et al.
Published: (2026)
by: Xie, Zhenghao, et al.
Published: (2026)
Homogeneous Tokenizer Matters: Homogeneous Visual Tokenizer for Remote Sensing Image Understanding
by: Shao, Run, et al.
Published: (2024)
by: Shao, Run, et al.
Published: (2024)
RSEdit: Text-Guided Image Editing for Remote Sensing
by: Zhenyuan, Chen, et al.
Published: (2026)
by: Zhenyuan, Chen, et al.
Published: (2026)
Task-Guided Prompting for Unified Remote Sensing Image Restoration
by: Huang, Wenli, et al.
Published: (2026)
by: Huang, Wenli, et al.
Published: (2026)
Aquila-plus: Prompt-Driven Visual-Language Models for Pixel-Level Remote Sensing Image Understanding
by: Lu, Kaixuan
Published: (2024)
by: Lu, Kaixuan
Published: (2024)
CSFMamba: Cross State Fusion Mamba Operator for Multimodal Remote Sensing Image Classification
by: Wang, Qingyu, et al.
Published: (2025)
by: Wang, Qingyu, et al.
Published: (2025)
MapGlue: Multimodal Remote Sensing Image Matching
by: Wu, Peihao, et al.
Published: (2025)
by: Wu, Peihao, et al.
Published: (2025)
Beyond Open Vocabulary: Multimodal Prompting for Object Detection in Remote Sensing Images
by: Yang, Shuai, et al.
Published: (2026)
by: Yang, Shuai, et al.
Published: (2026)
Balanced Diffusion-Guided Fusion for Multimodal Remote Sensing Classification
by: Liu, Hao, et al.
Published: (2025)
by: Liu, Hao, et al.
Published: (2025)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
by: Yang, Cong, et al.
Published: (2024)
by: Yang, Cong, et al.
Published: (2024)
CBEN -- A Multimodal Machine Learning Dataset for Cloud Robust Remote Sensing Image Understanding
by: Stricker, Marco, et al.
Published: (2026)
by: Stricker, Marco, et al.
Published: (2026)
Decoding the Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models
by: Li, Xiaohe, et al.
Published: (2026)
by: Li, Xiaohe, et al.
Published: (2026)
GRASP: Guided Region-Aware Sparse Prompting for Adapting MLLMs to Remote Sensing
by: Sun, Qigan, et al.
Published: (2026)
by: Sun, Qigan, et al.
Published: (2026)
Knowledge-aware Text-Image Retrieval for Remote Sensing Images
by: Mi, Li, et al.
Published: (2024)
by: Mi, Li, et al.
Published: (2024)
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering
by: Zhao, Zhicheng, et al.
Published: (2024)
by: Zhao, Zhicheng, et al.
Published: (2024)
MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
Multimodal-Aware Fusion Network for Referring Remote Sensing Image Segmentation
by: Shi, Leideng, et al.
Published: (2025)
by: Shi, Leideng, et al.
Published: (2025)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
by: Yuan, Bo, et al.
Published: (2024)
by: Yuan, Bo, et al.
Published: (2024)
Image-Conditioned Instance Prompt Network for Referring Remote Sensing Image Segmentation
by: Ren, Biaoyu, et al.
Published: (2026)
by: Ren, Biaoyu, et al.
Published: (2026)
Prompting DirectSAM for Semantic Contour Extraction in Remote Sensing Images
by: Miao, Shiyu, et al.
Published: (2024)
by: Miao, Shiyu, et al.
Published: (2024)
SGDiff: Scene Graph Guided Diffusion Model for Image Collaborative SegCaptioning
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Exploring Text-Guided Single Image Editing for Remote Sensing Images
by: Han, Fangzhou, et al.
Published: (2024)
by: Han, Fangzhou, et al.
Published: (2024)
SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning
by: Yang, Xiao, et al.
Published: (2026)
by: Yang, Xiao, et al.
Published: (2026)
IMDPrompter: Adapting SAM to Image Manipulation Detection by Cross-View Automated Prompt Learning
by: Zhang, Quan, et al.
Published: (2025)
by: Zhang, Quan, et al.
Published: (2025)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Context-Enhanced Detector For Building Detection From Remote Sensing Images
by: Huang, Ziyue, et al.
Published: (2023)
by: Huang, Ziyue, et al.
Published: (2023)
SA-MixNet: Structure-aware Mixup and Invariance Learning for Scribble-supervised Road Extraction in Remote Sensing Images
by: Feng, Jie, et al.
Published: (2024)
by: Feng, Jie, et al.
Published: (2024)
Towards Global Optimal Visual In-Context Learning Prompt Selection
by: Xu, Chengming, et al.
Published: (2024)
by: Xu, Chengming, et al.
Published: (2024)
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
by: Ou, Ruizhe, et al.
Published: (2025)
by: Ou, Ruizhe, et al.
Published: (2025)
RemoteDet-Mamba: A Hybrid Mamba-CNN Network for Multi-modal Object Detection in Remote Sensing Images
by: Ren, Kejun, et al.
Published: (2024)
by: Ren, Kejun, et al.
Published: (2024)
Cross-modal Prompting for Balanced Incomplete Multi-modal Emotion Recognition
by: He, Wen-Jue, et al.
Published: (2025)
by: He, Wen-Jue, et al.
Published: (2025)
Locality-aware Cross-modal Correspondence Learning for Dense Audio-Visual Events Localization
by: Xing, Ling, et al.
Published: (2024)
by: Xing, Ling, et al.
Published: (2024)
Similar Items
-
EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
by: Zhang, Wei, et al.
Published: (2024) -
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
by: Zhang, Xu, et al.
Published: (2026) -
PVPUFormer: Probabilistic Visual Prompt Unified Transformer for Interactive Image Segmentation
by: Zhang, Xu, et al.
Published: (2023) -
CtxMIM: Context-Enhanced Masked Image Modeling for Remote Sensing Image Understanding
by: Zhang, Mingming, et al.
Published: (2023) -
Learning a Cross-modality Anomaly Detector for Remote Sensing Imagery
by: Li, Jingtao, et al.
Published: (2023)