Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting
Fuente:
arXiv
Saved in:
| Main Authors: | Shah, Panav, Sethi, Geet, Gandhe, Ashutosh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery
by: Sethi, Geet, et al.
Published: (2026)
by: Sethi, Geet, et al.
Published: (2026)
GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
by: Zhou, Yue, et al.
Published: (2024)
by: Zhou, Yue, et al.
Published: (2024)
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
by: Li, Ke, et al.
Published: (2026)
by: Li, Ke, et al.
Published: (2026)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
by: Zhang, Peirong, et al.
Published: (2025)
by: Zhang, Peirong, et al.
Published: (2025)
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
by: Toker, Aysim, et al.
Published: (2025)
by: Toker, Aysim, et al.
Published: (2025)
Efficient Adaptation For Remote Sensing Visual Grounding
by: Moughnieh, Hasan, et al.
Published: (2025)
by: Moughnieh, Hasan, et al.
Published: (2025)
RemoteVAR: Autoregressive Visual Modeling for Remote Sensing Change Detection
by: Korkmaz, Yilmaz, et al.
Published: (2026)
by: Korkmaz, Yilmaz, et al.
Published: (2026)
RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning
by: Huang, Shiqi, et al.
Published: (2026)
by: Huang, Shiqi, et al.
Published: (2026)
Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models
by: Guo, Haonan, et al.
Published: (2024)
by: Guo, Haonan, et al.
Published: (2024)
ViRefSAM: Visual Reference-Guided Segment Anything Model for Remote Sensing Segmentation
by: Bi, Hanbo, et al.
Published: (2025)
by: Bi, Hanbo, et al.
Published: (2025)
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding
by: Zhu, Jiashun, et al.
Published: (2026)
by: Zhu, Jiashun, et al.
Published: (2026)
GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing
by: Shabbir, Akashah, et al.
Published: (2025)
by: Shabbir, Akashah, et al.
Published: (2025)
Improving Generalization in Visual Reasoning via Self-Ensemble
by: Nguyen, Tien-Huy, et al.
Published: (2024)
by: Nguyen, Tien-Huy, et al.
Published: (2024)
Contact-Aware Refinement of Human Pose Pseudo-Ground Truth via Bioimpedance Sensing
by: Forte, Maria-Paola, et al.
Published: (2025)
by: Forte, Maria-Paola, et al.
Published: (2025)
Text-Guided Coarse-to-Fine Fusion Network for Robust Remote Sensing Visual Question Answering
by: Zhao, Zhicheng, et al.
Published: (2024)
by: Zhao, Zhicheng, et al.
Published: (2024)
Text2Seg: Remote Sensing Image Semantic Segmentation via Text-Guided Visual Foundation Models
by: Zhang, Jielu, et al.
Published: (2023)
by: Zhang, Jielu, et al.
Published: (2023)
Employing Universal Voting Schemes for Improved Visual Place Recognition Performance
by: Waheed, Maria, et al.
Published: (2024)
by: Waheed, Maria, et al.
Published: (2024)
Cross-modal Context-aware Learning for Visual Prompt Guided Multimodal Image Understanding in Remote Sensing
by: Zhang, Xu, et al.
Published: (2025)
by: Zhang, Xu, et al.
Published: (2025)
RSEdit: Text-Guided Image Editing for Remote Sensing
by: Zhenyuan, Chen, et al.
Published: (2026)
by: Zhenyuan, Chen, et al.
Published: (2026)
SIGMAE: A Spectral-Index-Guided Foundation Model for Multispectral Remote Sensing
by: Zhang, Xiaokang, et al.
Published: (2026)
by: Zhang, Xiaokang, et al.
Published: (2026)
Referring Remote Sensing Image Segmentation via Bidirectional Alignment Guided Joint Prediction
by: Zhang, Tianxiang, et al.
Published: (2025)
by: Zhang, Tianxiang, et al.
Published: (2025)
DynamicVis: Dynamic Visual Perception for Efficient Remote Sensing Foundation Models
by: Chen, Keyan, et al.
Published: (2025)
by: Chen, Keyan, et al.
Published: (2025)
Semantic-Spatial Feature Fusion with Dynamic Graph Refinement for Remote Sensing Image Captioning
by: Liu, Maofu, et al.
Published: (2025)
by: Liu, Maofu, et al.
Published: (2025)
Language-Guided Diffusion Model for Visual Grounding
by: Chen, Sijia, et al.
Published: (2023)
by: Chen, Sijia, et al.
Published: (2023)
Improving Visual Reasoning with Iterative Evidence Refinement
by: Shi, Zeru, et al.
Published: (2026)
by: Shi, Zeru, et al.
Published: (2026)
OptiSAR-Net++: A Large-Scale Benchmark and Transformer-Free Framework for Cross-Domain Remote Sensing Visual Grounding
by: Tang, Xiaoyu, et al.
Published: (2026)
by: Tang, Xiaoyu, et al.
Published: (2026)
GeoMeld: Toward Semantically Grounded Foundation Models for Remote Sensing
by: Hasan, Maram, et al.
Published: (2026)
by: Hasan, Maram, et al.
Published: (2026)
Balanced Diffusion-Guided Fusion for Multimodal Remote Sensing Classification
by: Liu, Hao, et al.
Published: (2025)
by: Liu, Hao, et al.
Published: (2025)
PriorCLIP: Visual Prior Guided Vision-Language Model for Remote Sensing Image-Text Retrieval
by: Pan, Jiancheng, et al.
Published: (2024)
by: Pan, Jiancheng, et al.
Published: (2024)
RS-vHeat: Heat Conduction Guided Efficient Remote Sensing Foundation Model
by: Hu, Huiyang, et al.
Published: (2024)
by: Hu, Huiyang, et al.
Published: (2024)
Remote Sensing Image Classification Using Deep Ensemble Learning
by: Islam, Niful, et al.
Published: (2026)
by: Islam, Niful, et al.
Published: (2026)
Probabilistic Image-Driven Traffic Modeling via Remote Sensing
by: Workman, Scott, et al.
Published: (2024)
by: Workman, Scott, et al.
Published: (2024)
Large Vision-Language Models for Remote Sensing Visual Question Answering
by: Siripong, Surasakdi, et al.
Published: (2024)
by: Siripong, Surasakdi, et al.
Published: (2024)
Knowledge-aware Visual Question Generation for Remote Sensing Images
by: Li, Siran, et al.
Published: (2026)
by: Li, Siran, et al.
Published: (2026)
Visual Question Answering on Multiple Remote Sensing Image Modalities
by: Boussaid, Hichem, et al.
Published: (2025)
by: Boussaid, Hichem, et al.
Published: (2025)
Clustering Guided Domain-Specific Pretrained Foundation Model Very High-Resolution Arctic Remote Sensing
by: Perera, Amal S., et al.
Published: (2026)
by: Perera, Amal S., et al.
Published: (2026)
LR-FPN: Enhancing Remote Sensing Object Detection with Location Refined Feature Pyramid Network
by: Li, Hanqian, et al.
Published: (2024)
by: Li, Hanqian, et al.
Published: (2024)
Reconstruction Guided Few-shot Network For Remote Sensing Image Classification
by: Jaiswal, Mohit, et al.
Published: (2026)
by: Jaiswal, Mohit, et al.
Published: (2026)
Uncertainty-Guided Edge Learning for Deep Image Regression in Remote Sensing
by: Nguyen, Anh Vu, et al.
Published: (2026)
by: Nguyen, Anh Vu, et al.
Published: (2026)
Task-Guided Multi-Annotation Triplet Learning for Remote Sensing Representations
by: Zhou, Meilun, et al.
Published: (2026)
by: Zhou, Meilun, et al.
Published: (2026)
Similar Items
-
DiffuSAM: Diffusion Guided Zero-Shot Object Grounding for Remote Sensing Imagery
by: Sethi, Geet, et al.
Published: (2026) -
GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding
by: Zhou, Yue, et al.
Published: (2024) -
ProVG: Progressive Visual Grounding via Language Decoupling for Remote Sensing Imagery
by: Li, Ke, et al.
Published: (2026) -
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
by: Zhang, Peirong, et al.
Published: (2025) -
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
by: Toker, Aysim, et al.
Published: (2025)