EarthGPT: A Universal Multi-modal Large Language Model for Multi-sensor Image Comprehension in Remote Sensing Domain
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Wei, Cai, Miaoxin, Zhang, Tong, Zhuang, Yin, Mao, Xuerui |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
by: Jiang, Xiaozheng, et al.
Published: (2025)
by: Jiang, Xiaozheng, et al.
Published: (2025)
MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning
by: Wang, Tong, et al.
Published: (2025)
by: Wang, Tong, et al.
Published: (2025)
SegEarth-R2: Towards Comprehensive Language-guided Segmentation for Remote Sensing Images
by: Xin, Zepeng, et al.
Published: (2025)
by: Xin, Zepeng, et al.
Published: (2025)
GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation
by: Cai, Shihao, et al.
Published: (2024)
by: Cai, Shihao, et al.
Published: (2024)
GAIA: A Global, Multi-modal, Multi-scale Vision-Language Dataset for Remote Sensing Image Analysis
by: Zavras, Angelos, et al.
Published: (2025)
by: Zavras, Angelos, et al.
Published: (2025)
SMART-Ship: A Comprehensive Synchronized Multi-modal Aligned Remote Sensing Targets Dataset and Benchmark for Berthed Ships Analysis
by: Fan, Chen-Chen, et al.
Published: (2025)
by: Fan, Chen-Chen, et al.
Published: (2025)
SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery
by: Guo, Xin, et al.
Published: (2023)
by: Guo, Xin, et al.
Published: (2023)
JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation
by: Liu, Kai, et al.
Published: (2025)
by: Liu, Kai, et al.
Published: (2025)
SkySense V2: A Unified Foundation Model for Multi-modal Remote Sensing
by: Zhang, Yingying, et al.
Published: (2025)
by: Zhang, Yingying, et al.
Published: (2025)
GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
by: Ou, Ruizhe, et al.
Published: (2025)
by: Ou, Ruizhe, et al.
Published: (2025)
RingMoE: Mixture-of-Modality-Experts Multi-Modal Foundation Models for Universal Remote Sensing Image Interpretation
by: Bi, Hanbo, et al.
Published: (2025)
by: Bi, Hanbo, et al.
Published: (2025)
GroundingGPT:Language Enhanced Multi-modal Grounding Model
by: Li, Zhaowei, et al.
Published: (2024)
by: Li, Zhaowei, et al.
Published: (2024)
M$^3$amba: CLIP-driven Mamba Model for Multi-modal Remote Sensing Classification
by: Cao, Mingxiang, et al.
Published: (2025)
by: Cao, Mingxiang, et al.
Published: (2025)
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
by: Guo, Haonan, et al.
Published: (2024)
by: Guo, Haonan, et al.
Published: (2024)
Continual Panoptic Perception: Towards Multi-modal Incremental Interpretation of Remote Sensing Images
by: Yuan, Bo, et al.
Published: (2024)
by: Yuan, Bo, et al.
Published: (2024)
RemoteDet-Mamba: A Hybrid Mamba-CNN Network for Multi-modal Object Detection in Remote Sensing Images
by: Ren, Kejun, et al.
Published: (2024)
by: Ren, Kejun, et al.
Published: (2024)
Multi-modal Uncertainty Robust Tree Cover Segmentation For High-Resolution Remote Sensing Images
by: Gui, Yuanyuan, et al.
Published: (2025)
by: Gui, Yuanyuan, et al.
Published: (2025)
UniDet-D: A Unified Dynamic Spectral Attention Model for Object Detection under Adverse Weathers
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation
by: Gong, Ziyang, et al.
Published: (2024)
by: Gong, Ziyang, et al.
Published: (2024)
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
by: Wang, Yeyuan, et al.
Published: (2024)
by: Wang, Yeyuan, et al.
Published: (2024)
MMM-RS: A Multi-modal, Multi-GSD, Multi-scene Remote Sensing Dataset and Benchmark for Text-to-Image Generation
by: Luo, Jialin, et al.
Published: (2024)
by: Luo, Jialin, et al.
Published: (2024)
DescribeEarth: Describe Anything for Remote Sensing Images
by: Li, Kaiyu, et al.
Published: (2025)
by: Li, Kaiyu, et al.
Published: (2025)
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
by: Yang, Rui, et al.
Published: (2025)
by: Yang, Rui, et al.
Published: (2025)
Large Multi-modality Model Assisted AI-Generated Image Quality Assessment
by: Wang, Puyi, et al.
Published: (2024)
by: Wang, Puyi, et al.
Published: (2024)
Coherent and Multi-modality Image Inpainting via Latent Space Optimization
by: Pan, Lingzhi, et al.
Published: (2024)
by: Pan, Lingzhi, et al.
Published: (2024)
PhotoFramer: Multi-modal Image Composition Instruction
by: You, Zhiyuan, et al.
Published: (2025)
by: You, Zhiyuan, et al.
Published: (2025)
Attribution-Guided Multimodal Deepfake Detection via Cross-Modal Forensic Fingerprints
by: Ahmad, Wasim, et al.
Published: (2026)
by: Ahmad, Wasim, et al.
Published: (2026)
Multi-level Cross-modal Alignment for Image Clustering
by: Qiu, Liping, et al.
Published: (2024)
by: Qiu, Liping, et al.
Published: (2024)
Multimodal Interpretation of Remote Sensing Images: Dynamic Resolution Input Strategy and Multi-scale Vision-Language Alignment Mechanism
by: Zhang, Siyu, et al.
Published: (2025)
by: Zhang, Siyu, et al.
Published: (2025)
DDLNet: Boosting Remote Sensing Change Detection with Dual-Domain Learning
by: Ma, Xiaowen, et al.
Published: (2024)
by: Ma, Xiaowen, et al.
Published: (2024)
Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation
by: Liu, Sihan, et al.
Published: (2023)
by: Liu, Sihan, et al.
Published: (2023)
P-MSDiff: Parallel Multi-Scale Diffusion for Remote Sensing Image Segmentation
by: Zhang, Qi, et al.
Published: (2024)
by: Zhang, Qi, et al.
Published: (2024)
RingMo-Agent: A Unified Remote Sensing Foundation Model for Multi-Platform and Multi-Modal Reasoning
by: Hu, Huiyang, et al.
Published: (2025)
by: Hu, Huiyang, et al.
Published: (2025)
EarthBridge: A Solution for 4th Multi-modal Aerial View Image Challenge Translation Track
by: Chen, Zhenyuan, et al.
Published: (2026)
by: Chen, Zhenyuan, et al.
Published: (2026)
ESC-MISR: Enhancing Spatial Correlations for Multi-Image Super-Resolution in Remote Sensing
by: Zhang, Zhihui, et al.
Published: (2024)
by: Zhang, Zhihui, et al.
Published: (2024)
AlignGPT: Multi-modal Large Language Models with Adaptive Alignment Capability
by: Zhao, Fei, et al.
Published: (2024)
by: Zhao, Fei, et al.
Published: (2024)
Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
Similar Items
-
EarthGPT-X: A Spatial MLLM for Multi-level Multi-Source Remote Sensing Imagery Understanding with Visual Prompting
by: Zhang, Wei, et al.
Published: (2025) -
EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
by: Zhang, Wei, et al.
Published: (2024) -
Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery
by: Zhang, Wei, et al.
Published: (2024) -
RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
by: Jiang, Xiaozheng, et al.
Published: (2025) -
MSSDF: Modality-Shared Self-supervised Distillation for High-Resolution Multi-modal Remote Sensing Image Learning
by: Wang, Tong, et al.
Published: (2025)