Efficient Adaptation For Remote Sensing Visual Grounding

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Moughnieh, Hasan, Chalhoub, Mohamad, Nasrallah, Hasan, Nattero, Cristiano, Campanella, Paolo, Nico, Giovanni, Ghandour, Ali J.
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915313453367296
author Moughnieh, Hasan
Chalhoub, Mohamad
Nasrallah, Hasan
Nattero, Cristiano
Campanella, Paolo
Nico, Giovanni
Ghandour, Ali J.
author_facet Moughnieh, Hasan
Chalhoub, Mohamad
Nasrallah, Hasan
Nattero, Cristiano
Campanella, Paolo
Nico, Giovanni
Ghandour, Ali J.
contents Adapting pre-trained models has become an effective strategy in artificial intelligence, offering a scalable and efficient alternative to training models from scratch. In the context of remote sensing (RS), where visual grounding(VG) remains underexplored, this approach enables the deployment of powerful vision-language models to achieve robust cross-modal understanding while significantly reducing computational overhead. To address this, we applied Parameter Efficient Fine Tuning (PEFT) techniques to adapt these models for RS-specific VG tasks. Specifically, we evaluated LoRA placement across different modules in Grounding DINO and used BitFit and adapters to fine-tune the OFA foundation model pre-trained on general-purpose VG datasets. This approach achieved performance comparable to or surpassing current State Of The Art (SOTA) models while significantly reducing computational costs. This study highlights the potential of PEFT techniques to advance efficient and precise multi-modal analysis in RS, offering a practical and cost-effective alternative to full model training.
format Preprint
id arxiv_https___arxiv_org_abs_2503_23083
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Efficient Adaptation For Remote Sensing Visual Grounding
Moughnieh, Hasan
Chalhoub, Mohamad
Nasrallah, Hasan
Nattero, Cristiano
Campanella, Paolo
Nico, Giovanni
Ghandour, Ali J.
Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
Adapting pre-trained models has become an effective strategy in artificial intelligence, offering a scalable and efficient alternative to training models from scratch. In the context of remote sensing (RS), where visual grounding(VG) remains underexplored, this approach enables the deployment of powerful vision-language models to achieve robust cross-modal understanding while significantly reducing computational overhead. To address this, we applied Parameter Efficient Fine Tuning (PEFT) techniques to adapt these models for RS-specific VG tasks. Specifically, we evaluated LoRA placement across different modules in Grounding DINO and used BitFit and adapters to fine-tune the OFA foundation model pre-trained on general-purpose VG datasets. This approach achieved performance comparable to or surpassing current State Of The Art (SOTA) models while significantly reducing computational costs. This study highlights the potential of PEFT techniques to advance efficient and precise multi-modal analysis in RS, offering a practical and cost-effective alternative to full model training.
title Efficient Adaptation For Remote Sensing Visual Grounding
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Computation and Language
url https://arxiv.org/abs/2503.23083