DGTRSD & DGTRS-CLIP: A Dual-Granularity Remote Sensing Image-Text Dataset and Vision Language Foundation Model for Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Weizhi, Deng, Yupeng, Wei, Jin, Chen, Jingbo, Chen, Jiansheng, Feng, Yuman, Xi, Zhihao, Liu, Diyou, Li, Kai, Meng, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GLD-Road:A global-local decoding road network extraction model for remote sensing images
by: Deng, Ligao, et al.
Published: (2025)
by: Deng, Ligao, et al.
Published: (2025)
IRSAMap:Towards Large-Scale, High-Resolution Land Cover Map Vectorization
by: Meng, Yu, et al.
Published: (2025)
by: Meng, Yu, et al.
Published: (2025)
PolyFootNet: Extracting Polygonal Building Footprints in Off-Nadir Remote Sensing Images
by: Li, Kai, et al.
Published: (2024)
by: Li, Kai, et al.
Published: (2024)
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
by: Liu, Fan, et al.
Published: (2023)
by: Liu, Fan, et al.
Published: (2023)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
by: Zohra, Fatimah, et al.
Published: (2025)
by: Zohra, Fatimah, et al.
Published: (2025)
Prompt-Driven Building Footprint Extraction in Aerial Images with Offset-Building Model
by: Li, Kai, et al.
Published: (2023)
by: Li, Kai, et al.
Published: (2023)
GeoAlignCLIP: Enhancing Fine-Grained Vision-Language Alignment in Remote Sensing via Multi-Granular Consistency Learning
by: Yang, Xiao, et al.
Published: (2026)
by: Yang, Xiao, et al.
Published: (2026)
PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval
by: Pan, Jiancheng, et al.
Published: (2024)
by: Pan, Jiancheng, et al.
Published: (2024)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
by: Wei, Zhixiang, et al.
Published: (2025)
by: Wei, Zhixiang, et al.
Published: (2025)
CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation
by: Gong, Ziyang, et al.
Published: (2024)
by: Gong, Ziyang, et al.
Published: (2024)
Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing
by: Chen, Xi, et al.
Published: (2026)
by: Chen, Xi, et al.
Published: (2026)
Text2Earth: Unlocking Text-driven Remote Sensing Image Generation with a Global-Scale Dataset and a Foundation Model
by: Liu, Chenyang, et al.
Published: (2025)
by: Liu, Chenyang, et al.
Published: (2025)
FLAVARS: A Multimodal Foundational Language and Vision Alignment Model for Remote Sensing
by: Corley, Isaac, et al.
Published: (2025)
by: Corley, Isaac, et al.
Published: (2025)
MMLGNet: Cross-Modal Alignment of Remote Sensing Data using CLIP
by: Chaudhary, Aditya, et al.
Published: (2026)
by: Chaudhary, Aditya, et al.
Published: (2026)
Adapting Vision Foundation Models for Robust Cloud Segmentation in Remote Sensing Images
by: Zou, Xuechao, et al.
Published: (2024)
by: Zou, Xuechao, et al.
Published: (2024)
PeftCD: Leveraging Vision Foundation Models with Parameter-Efficient Fine-Tuning for Remote Sensing Change Detection
by: Dong, Sijun, et al.
Published: (2025)
by: Dong, Sijun, et al.
Published: (2025)
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
by: Wang, Junjue, et al.
Published: (2025)
by: Wang, Junjue, et al.
Published: (2025)
ChangeAnywhere: Sample Generation for Remote Sensing Change Detection via Semantic Latent Diffusion Model
by: Tang, Kai, et al.
Published: (2024)
by: Tang, Kai, et al.
Published: (2024)
Anti‐Swelling Textile Power Generator with 1D Nanoscale Channel Alignment in Nanofiber/Graphene Hybrid Yarns
by: Yuman Zhou, et al.
Published: (2025)
by: Yuman Zhou, et al.
Published: (2025)
Dual-Branch Remote Sensing Infrared Image Super-Resolution
by: Ge, Xining, et al.
Published: (2026)
by: Ge, Xining, et al.
Published: (2026)
CIBR: Cross-modal Information Bottleneck Regularization for Robust CLIP Generalization
by: Ji, Yingrui, et al.
Published: (2025)
by: Ji, Yingrui, et al.
Published: (2025)
A General Adaptive Dual-level Weighting Mechanism for Remote Sensing Pansharpening
by: Huang, Jie, et al.
Published: (2025)
by: Huang, Jie, et al.
Published: (2025)
Vision Foundation Models in Remote Sensing: A Survey
by: Lu, Siqi, et al.
Published: (2024)
by: Lu, Siqi, et al.
Published: (2024)
Multi-Perspective Subimage CLIP with Keyword Guidance for Remote Sensing Image-Text Retrieval
by: Li, Yifan, et al.
Published: (2026)
by: Li, Yifan, et al.
Published: (2026)
RSEdit: Text-Guided Image Editing for Remote Sensing
by: Zhenyuan, Chen, et al.
Published: (2026)
by: Zhenyuan, Chen, et al.
Published: (2026)
MoralCLIP: Contrastive Alignment of Vision-and-Language Representations with Moral Foundations Theory
by: Condez, Ana Carolina, et al.
Published: (2025)
by: Condez, Ana Carolina, et al.
Published: (2025)
TimeSenCLIP: A Time Series Vision-Language Model for Remote Sensing
by: Jain, Pallavi, et al.
Published: (2025)
by: Jain, Pallavi, et al.
Published: (2025)
Fourier Angle Alignment for Oriented Object Detection in Remote Sensing
by: Gu, Changyu, et al.
Published: (2026)
by: Gu, Changyu, et al.
Published: (2026)
STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning
by: Gong, Yanpei, et al.
Published: (2026)
by: Gong, Yanpei, et al.
Published: (2026)
HiMo-CLIP: Modeling Semantic Hierarchy and Monotonicity in Vision-Language Alignment
by: Wu, Ruijia, et al.
Published: (2025)
by: Wu, Ruijia, et al.
Published: (2025)
Falcon: A Remote Sensing Vision-Language Foundation Model (Technical Report)
by: Yao, Kelu, et al.
Published: (2025)
by: Yao, Kelu, et al.
Published: (2025)
SmartCLIP: Modular Vision-language Alignment with Identification Guarantees
by: Xie, Shaoan, et al.
Published: (2025)
by: Xie, Shaoan, et al.
Published: (2025)
Defect Spectrum: A Granular Look of Large-Scale Defect Datasets with Rich Semantics
by: Yang, Shuai, et al.
Published: (2023)
by: Yang, Shuai, et al.
Published: (2023)
Dataset: Application of Remote Sensing and Machine Learning Algorithms for Shipwreck Susceptibility Mapping in China
by: Chen, Junhui
Published: (2025)
by: Chen, Junhui
Published: (2025)
RSDehamba: Lightweight Vision Mamba for Remote Sensing Satellite Image Dehazing
by: Zhou, Huiling, et al.
Published: (2024)
by: Zhou, Huiling, et al.
Published: (2024)
VFM-ISRefiner: Towards Better Adapting Vision Foundation Models for Interactive Segmentation of Remote Sensing Images
by: Wang, Deliang, et al.
Published: (2025)
by: Wang, Deliang, et al.
Published: (2025)
UMG-CLIP: A Unified Multi-Granularity Vision Generalist for Open-World Understanding
by: Shi, Bowen, et al.
Published: (2024)
by: Shi, Bowen, et al.
Published: (2024)
DynamicVis: Dynamic Visual Perception for Efficient Remote Sensing Foundation Models
by: Chen, Keyan, et al.
Published: (2025)
by: Chen, Keyan, et al.
Published: (2025)
Probabilistic Formulations for System Identification of Linear Dynamics with Bilinear Observation Models
by: Liu, Diyou, et al.
Published: (2025)
by: Liu, Diyou, et al.
Published: (2025)
MER-CLIP: AU-Guided Vision-Language Alignment for Micro-Expression Recognition
by: Liu, Shifeng, et al.
Published: (2025)
by: Liu, Shifeng, et al.
Published: (2025)
Similar Items
-
GLD-Road:A global-local decoding road network extraction model for remote sensing images
by: Deng, Ligao, et al.
Published: (2025) -
IRSAMap:Towards Large-Scale, High-Resolution Land Cover Map Vectorization
by: Meng, Yu, et al.
Published: (2025) -
PolyFootNet: Extracting Polygonal Building Footprints in Off-Nadir Remote Sensing Images
by: Li, Kai, et al.
Published: (2024) -
RemoteCLIP: A Vision Language Foundation Model for Remote Sensing
by: Liu, Fan, et al.
Published: (2023) -
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
by: Zohra, Fatimah, et al.
Published: (2025)