Towards Efficient Benchmarking of Foundation Models in Remote Sensing: A Capabilities Encoding Approach

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Adorni, Pierre, Pham, Minh-Tan, May, Stéphane, Lefèvre, Sébastien
Format: Preprint
Published: 2025
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866912362199515136
author Adorni, Pierre
Pham, Minh-Tan
May, Stéphane
Lefèvre, Sébastien
author_facet Adorni, Pierre
Pham, Minh-Tan
May, Stéphane
Lefèvre, Sébastien
contents Foundation models constitute a significant advancement in computer vision: after a single, albeit costly, training phase, they can address a wide array of tasks. In the field of Earth observation, over 75 remote sensing vision foundation models have been developed in the past four years. However, none has consistently outperformed the others across all available downstream tasks. To facilitate their comparison, we propose a cost-effective method for predicting a model's performance on multiple downstream tasks without the need for fine-tuning on each one. This method is based on what we call "capabilities encoding." The utility of this novel approach is twofold: we demonstrate its potential to simplify the selection of a foundation model for a given new task, and we employ it to offer a fresh perspective on the existing literature, suggesting avenues for future research. Codes are available at https://github.com/pierreadorni/capabilities-encoding.
format Preprint
id arxiv_https___arxiv_org_abs_2505_03299
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Towards Efficient Benchmarking of Foundation Models in Remote Sensing: A Capabilities Encoding Approach
Adorni, Pierre
Pham, Minh-Tan
May, Stéphane
Lefèvre, Sébastien
Computer Vision and Pattern Recognition
Artificial Intelligence
Foundation models constitute a significant advancement in computer vision: after a single, albeit costly, training phase, they can address a wide array of tasks. In the field of Earth observation, over 75 remote sensing vision foundation models have been developed in the past four years. However, none has consistently outperformed the others across all available downstream tasks. To facilitate their comparison, we propose a cost-effective method for predicting a model's performance on multiple downstream tasks without the need for fine-tuning on each one. This method is based on what we call "capabilities encoding." The utility of this novel approach is twofold: we demonstrate its potential to simplify the selection of a foundation model for a given new task, and we employ it to offer a fresh perspective on the existing literature, suggesting avenues for future research. Codes are available at https://github.com/pierreadorni/capabilities-encoding.
title Towards Efficient Benchmarking of Foundation Models in Remote Sensing: A Capabilities Encoding Approach
topic Computer Vision and Pattern Recognition
Artificial Intelligence
url https://arxiv.org/abs/2505.03299