GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Peirong, Zhang, Yidan, Xu, Luxiao, Lin, Jinliang, Guo, Zonghao, Wang, Fengxiang, Yang, Xue, Wei, Kaiwen, Wang, Lei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
by: Wang, Fengxiang, et al.
Published: (2026)
by: Wang, Fengxiang, et al.
Published: (2026)
GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
by: Zhu, Jiashun, et al.
Published: (2026)
by: Zhu, Jiashun, et al.
Published: (2026)
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
Efficient Adaptation For Remote Sensing Visual Grounding
by: Moughnieh, Hasan, et al.
Published: (2025)
by: Moughnieh, Hasan, et al.
Published: (2025)
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
by: Bi, Hanbo, et al.
Published: (2025)
by: Bi, Hanbo, et al.
Published: (2025)
XLRS-Bench: Could Your Multimodal LLMs Understand Extremely Large Ultra-High-Resolution Remote Sensing Imagery?
by: Wang, Fengxiang, et al.
Published: (2025)
by: Wang, Fengxiang, et al.
Published: (2025)
Semantic-Geometric Dual Compression: Training-Free Visual Token Reduction for Ultra-High-Resolution Remote Sensing Understanding
by: Li, Yueying, et al.
Published: (2026)
by: Li, Yueying, et al.
Published: (2026)
Geospatial-Reasoning-Driven Vocabulary-Agnostic Remote Sensing Semantic Segmentation
by: Zhou, Chufeng, et al.
Published: (2026)
by: Zhou, Chufeng, et al.
Published: (2026)
AlignCAT: Visual-Linguistic Alignment of Category and Attribute for Weakly Supervised Visual Grounding
by: Wang, Yidan, et al.
Published: (2025)
by: Wang, Yidan, et al.
Published: (2025)
GeoZero: Incentivizing Reasoning from Scratch on Geospatial Scenes
by: Wang, Di, et al.
Published: (2025)
by: Wang, Di, et al.
Published: (2025)
ViCor: Bridging Visual Understanding and Commonsense Reasoning with Large Language Models
by: Zhou, Kaiwen, et al.
Published: (2023)
by: Zhou, Kaiwen, et al.
Published: (2023)
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
by: Guo, Haonan, et al.
Published: (2024)
by: Guo, Haonan, et al.
Published: (2024)
RAPTOR: Real-Time High-Resolution UAV Video Prediction with Efficient Video Attention
by: Chen, Zhan, et al.
Published: (2025)
by: Chen, Zhan, et al.
Published: (2025)
RSVG-ZeroOV: Exploring a Training-Free Framework for Zero-Shot Open-Vocabulary Visual Grounding in Remote Sensing Images
by: Li, Ke, et al.
Published: (2025)
by: Li, Ke, et al.
Published: (2025)
Advancements in Visual Language Models for Remote Sensing: Datasets, Capabilities, and Enhancement Techniques
by: Tao, Lijie, et al.
Published: (2024)
by: Tao, Lijie, et al.
Published: (2024)
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
by: Toker, Aysim, et al.
Published: (2025)
by: Toker, Aysim, et al.
Published: (2025)
The relationship between Chinese citizens' debt and redistribution preferences
by: Luxiao Wang, et al.
Published: (2025)
by: Luxiao Wang, et al.
Published: (2025)
CrossEarth: Geospatial Vision Foundation Model for Domain Generalizable Remote Sensing Semantic Segmentation
by: Gong, Ziyang, et al.
Published: (2024)
by: Gong, Ziyang, et al.
Published: (2024)
ViDA: Homeostatic Visual Domain Adapter for Continual Test Time Adaptation
by: Liu, Jiaming, et al.
Published: (2023)
by: Liu, Jiaming, et al.
Published: (2023)
LLaVA-UHD v3: Progressive Visual Compression for Efficient Native-Resolution Encoding in MLLMs
by: Sun, Shichu, et al.
Published: (2025)
by: Sun, Shichu, et al.
Published: (2025)
Deep Semantic-Visual Alignment for Zero-Shot Remote Sensing Image Scene Classification
by: Xu, Wenjia, et al.
Published: (2024)
by: Xu, Wenjia, et al.
Published: (2024)
ViGoR: Improving Visual Grounding of Large Vision Language Models with Fine-Grained Reward Modeling
by: Yan, Siming, et al.
Published: (2024)
by: Yan, Siming, et al.
Published: (2024)
RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
by: Feng, Sicheng, et al.
Published: (2025)
by: Feng, Sicheng, et al.
Published: (2025)
Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery
by: Zhang, Wei, et al.
Published: (2024)
by: Zhang, Wei, et al.
Published: (2024)
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
by: Xu, Xiaoxu, et al.
Published: (2023)
by: Xu, Xiaoxu, et al.
Published: (2023)
LeMeViT: Efficient Vision Transformer with Learnable Meta Tokens for Remote Sensing Image Interpretation
by: Jiang, Wentao, et al.
Published: (2024)
by: Jiang, Wentao, et al.
Published: (2024)
Multi-Agent Geospatial Copilots for Remote Sensing Workflows
by: Lee, Chaehong, et al.
Published: (2025)
by: Lee, Chaehong, et al.
Published: (2025)
AtrousMamaba: An Atrous-Window Scanning Visual State Space Model for Remote Sensing Change Detection
by: Wang, Tao, et al.
Published: (2025)
by: Wang, Tao, et al.
Published: (2025)
RewardDance: Reward Scaling in Visual Generation
by: Wu, Jie, et al.
Published: (2025)
by: Wu, Jie, et al.
Published: (2025)
MB-ORES: A Multi-Branch Object Reasoner for Visual Grounding in Remote Sensing
by: Radouane, Karim, et al.
Published: (2025)
by: Radouane, Karim, et al.
Published: (2025)
RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
LWGANet: Addressing Spatial and Channel Redundancy in Remote Sensing Visual Tasks with Light-Weight Grouped Attention
by: Lu, Wei, et al.
Published: (2025)
by: Lu, Wei, et al.
Published: (2025)
ViG-Bias: Visually Grounded Bias Discovery and Mitigation
by: Marani, Badr-Eddine, et al.
Published: (2024)
by: Marani, Badr-Eddine, et al.
Published: (2024)
Route Packing: Geospatially-Accurate Visualization of Route Networks
by: Zhao, Jieqiong, et al.
Published: (2019)
by: Zhao, Jieqiong, et al.
Published: (2019)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
by: Zhang, Jiaxi, et al.
Published: (2026)
by: Zhang, Jiaxi, et al.
Published: (2026)
MEET: A Million-Scale Dataset for Fine-Grained Geospatial Scene Classification with Zoom-Free Remote Sensing Imagery
by: Li, Yansheng, et al.
Published: (2025)
by: Li, Yansheng, et al.
Published: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
by: Wang, Yuduo, et al.
Published: (2023)
by: Wang, Yuduo, et al.
Published: (2023)
SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning
by: Yang, Xiao, et al.
Published: (2026)
by: Yang, Xiao, et al.
Published: (2026)
Similar Items
-
GeoEyes: On-Demand Visual Focusing for Evidence-Grounded Understanding of Ultra-High-Resolution Remote Sensing Imagery
by: Wang, Fengxiang, et al.
Published: (2026) -
GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
by: Zhou, Yue, et al.
Published: (2024) -
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
by: Zhu, Jiashun, et al.
Published: (2026) -
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
by: Li, Ke, et al.
Published: (2026) -
GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution
by: Wang, Fengxiang, et al.
Published: (2025)