TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation
Fuente:
arXiv
Saved in:
| Main Authors: | Shu, Yan, Ren, Bin, Xiong, Zhitong, Zhu, Xiao Xiang, Demir, Begüm, Sebe, Nicu, Rota, Paolo |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
by: Herzog, Johann-Ludwig, et al.
Published: (2026)
by: Herzog, Johann-Ludwig, et al.
Published: (2026)
EarthNets: Empowering AI in Earth Observation
by: Xiong, Zhitong, et al.
Published: (2022)
by: Xiong, Zhitong, et al.
Published: (2022)
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025)
by: Liu, Jiaqi, et al.
Published: (2025)
Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery
by: Burgert, Tom, et al.
Published: (2026)
by: Burgert, Tom, et al.
Published: (2026)
ImageNet-trained CNNs are not biased towards texture: Revisiting feature reliance through controlled suppression
by: Burgert, Tom, et al.
Published: (2025)
by: Burgert, Tom, et al.
Published: (2025)
One for All: Toward Unified Foundation Models for Earth Vision
by: Xiong, Zhitong, et al.
Published: (2024)
by: Xiong, Zhitong, et al.
Published: (2024)
Beyond Grid Data: Exploring Graph Neural Networks for Earth Observation
by: Zhao, Shan, et al.
Published: (2024)
by: Zhao, Shan, et al.
Published: (2024)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
EO-VAE: Towards A Multi-sensor Tokenizer for Earth Observation Data
by: Lehmann, Nils, et al.
Published: (2026)
by: Lehmann, Nils, et al.
Published: (2026)
Visual Text Processing: A Comprehensive Review and Unified Evaluation
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Hierarchical Cross-Attention Network for Virtual Try-On
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
REOBench: Benchmarking Robustness of Earth Observation Foundation Models
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
by: Yuan, Zhenghang, et al.
Published: (2024)
by: Yuan, Zhenghang, et al.
Published: (2024)
Vision+X: A Survey on Multimodal Learning in the Light of Data
by: Zhu, Ye, et al.
Published: (2022)
by: Zhu, Ye, et al.
Published: (2022)
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2025)
by: Xiong, Zhitong, et al.
Published: (2025)
Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation
by: Xiong, Zhitong, et al.
Published: (2024)
by: Xiong, Zhitong, et al.
Published: (2024)
Asymmetric GANs for Image-to-Image Translation
by: Tang, Hao, et al.
Published: (2019)
by: Tang, Hao, et al.
Published: (2019)
UrbanSARFloods: Sentinel-1 SLC-Based Benchmark Dataset for Urban and Open-Area Flood Mapping
by: Zhao, Jie, et al.
Published: (2024)
by: Zhao, Jie, et al.
Published: (2024)
TerraMesh: A Planetary Mosaic of Multimodal Earth Observation Data
by: Blumenstiel, Benedikt, et al.
Published: (2025)
by: Blumenstiel, Benedikt, et al.
Published: (2025)
TerraMind: Large-Scale Generative Multimodality for Earth Observation
by: Jakubik, Johannes, et al.
Published: (2025)
by: Jakubik, Johannes, et al.
Published: (2025)
CLIP is Strong Enough to Fight Back: Test-time Counterattacks towards Zero-shot Adversarial Robustness of CLIP
by: Xing, Songlong, et al.
Published: (2025)
by: Xing, Songlong, et al.
Published: (2025)
TerraCodec: Compressing Optical Earth Observation Data
by: Costa-Watanabe, Julen, et al.
Published: (2025)
by: Costa-Watanabe, Julen, et al.
Published: (2025)
Loomis Painter: Reconstructing the Painting Process
by: Pobitzer, Markus, et al.
Published: (2025)
by: Pobitzer, Markus, et al.
Published: (2025)
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
by: Lobba, Davide, et al.
Published: (2025)
by: Lobba, Davide, et al.
Published: (2025)
Dress-ED: Instruction-Guided Editing for Virtual Try-On and Try-Off
by: Sanguigni, Fulvio, et al.
Published: (2026)
by: Sanguigni, Fulvio, et al.
Published: (2026)
MagicBathyNet: A Multimodal Remote Sensing Dataset for Bathymetry Prediction and Pixel-based Classification in Shallow Waters
by: Agrafiotis, Panagiotis, et al.
Published: (2024)
by: Agrafiotis, Panagiotis, et al.
Published: (2024)
RankFeat&RankWeight: Rank-1 Feature/Weight Removal for Out-of-distribution Detection
by: Song, Yue, et al.
Published: (2023)
by: Song, Yue, et al.
Published: (2023)
Textual Knowledge Matters: Cross-Modality Co-Teaching for Generalized Visual Class Discovery
by: Zheng, Haiyang, et al.
Published: (2024)
by: Zheng, Haiyang, et al.
Published: (2024)
Hierarchical Semi-Supervised Active Learning for Remote Sensing
by: Huang, Wei, et al.
Published: (2025)
by: Huang, Wei, et al.
Published: (2025)
GAIA: A Global, Multi-modal, Multi-scale Vision-Language Dataset for Remote Sensing Image Analysis
by: Zavras, Angelos, et al.
Published: (2025)
by: Zavras, Angelos, et al.
Published: (2025)
Global-Local Distillation Network-Based Audio-Visual Speaker Tracking with Incomplete Modalities
by: Li, Yidi, et al.
Published: (2024)
by: Li, Yidi, et al.
Published: (2024)
Towards Unified Vision Language Models for Forest Ecological Analysis in Earth Observation
by: Xue, Xizhe, et al.
Published: (2025)
by: Xue, Xizhe, et al.
Published: (2025)
Hyperbolic Busemann Neural Networks
by: Chen, Ziheng, et al.
Published: (2026)
by: Chen, Ziheng, et al.
Published: (2026)
Reverse Personalization
by: Kung, Han-Wei, et al.
Published: (2025)
by: Kung, Han-Wei, et al.
Published: (2025)
TerraFlow: Multimodal, Multitemporal Representation Learning for Earth Observation
by: Puriy, Nazar, et al.
Published: (2026)
by: Puriy, Nazar, et al.
Published: (2026)
Democratizing Fine-grained Visual Recognition with Large Language Models
by: Liu, Mingxuan, et al.
Published: (2024)
by: Liu, Mingxuan, et al.
Published: (2024)
Masked Clustering Prediction for Unsupervised Point Cloud Pre-training
by: Ren, Bin, et al.
Published: (2025)
by: Ren, Bin, et al.
Published: (2025)
DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
by: Wu, Yihong, et al.
Published: (2025)
by: Wu, Yihong, et al.
Published: (2025)
ZeroReg: Zero-Shot Point Cloud Registration with Foundation Models
by: Wang, Weijie, et al.
Published: (2023)
by: Wang, Weijie, et al.
Published: (2023)
Similar Items
-
EarthMind: Leveraging Cross-Sensor Data for Advanced Earth Observation Interpretation with a Unified Multimodal LLM
by: Shu, Yan, et al.
Published: (2025) -
BigEarthNet.txt: A Large-Scale Multi-Sensor Image-Text Dataset and Benchmark for Earth Observation
by: Herzog, Johann-Ludwig, et al.
Published: (2026) -
EarthNets: Empowering AI in Earth Observation
by: Xiong, Zhitong, et al.
Published: (2022) -
Multi-focal Conditioned Latent Diffusion for Person Image Synthesis
by: Liu, Jiaqi, et al.
Published: (2025) -
Rank-based Geographical Regularization: Revisiting Contrastive Self-Supervised Learning for Multispectral Remote Sensing Imagery
by: Burgert, Tom, et al.
Published: (2026)