Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Shah, Panav, Sethi, Geet, Gandhe, Ashutosh
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914620313174016
author Shah, Panav
Sethi, Geet
Gandhe, Ashutosh
author_facet Shah, Panav
Sethi, Geet
Gandhe, Ashutosh
contents Visual grounding aims to locate image regions that correspond to natural language descriptions and is a key component of interpretable vision systems. In remote sensing imagery, grounding is particularly challenging due to complex scenes, small objects, and large variations in scale. Relying on a single model is often insufficient to address these diverse challenges. In this work, we propose two grounding pipelines, Sequential Grounding Refinement (SGR) and Cluster-Aware Grounding Refinement (CGR), that combine the complementary strengths of RemoteSAM, a visual grounding model specialized for remote sensing, and SAM3, a powerful general-purpose segmentation model. Our approach first uses RemoteSAM to obtain an initial estimate of object location, which is then refined using SAM3 to produce more accurate and spatially consistent segmentations. Additionally, we explore an ensemble strategy based on majority voting across six diverse grounding pipelines, each with distinct capabilities. This multi-model framework improves robustness and significantly enhances localization accuracy. Experimental results demonstrate that the proposed pipelines and ensemble approach outperform individual models, leading to more reliable and precise visual grounding predictions.
format Preprint
id arxiv_https___arxiv_org_abs_2606_00556
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting
Shah, Panav
Sethi, Geet
Gandhe, Ashutosh
Computer Vision and Pattern Recognition
Visual grounding aims to locate image regions that correspond to natural language descriptions and is a key component of interpretable vision systems. In remote sensing imagery, grounding is particularly challenging due to complex scenes, small objects, and large variations in scale. Relying on a single model is often insufficient to address these diverse challenges. In this work, we propose two grounding pipelines, Sequential Grounding Refinement (SGR) and Cluster-Aware Grounding Refinement (CGR), that combine the complementary strengths of RemoteSAM, a visual grounding model specialized for remote sensing, and SAM3, a powerful general-purpose segmentation model. Our approach first uses RemoteSAM to obtain an initial estimate of object location, which is then refined using SAM3 to produce more accurate and spatially consistent segmentations. Additionally, we explore an ensemble strategy based on majority voting across six diverse grounding pipelines, each with distinct capabilities. This multi-model framework improves robustness and significantly enhances localization accuracy. Experimental results demonstrate that the proposed pipelines and ensemble approach outperform individual models, leading to more reliable and precise visual grounding predictions.
title Improving Visual Grounding in Remote Sensing via Cluster-Guided Refinement and Model Ensemble Voting
topic Computer Vision and Pattern Recognition
url https://arxiv.org/abs/2606.00556