Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards

Fuente: arXiv
Guardado en:
Detalles Bibliográficos
Autores principales: Xu, Chenhui, Yu, Fuxun, Bianco, Michael J., Kovarskiy, Jacob, Tang, Raphael, Zhang, Qi, Xu, Zirui, LeVine, Will, Dubbs, Brandon, Liao, Heming, Burgess, Cassandra, Bag, Suvam, Patravali, Jay, Kukal, Rupanjali, Figueroa, Mikael, Madhok, Rishi, Karianakis, Nikolaos, Xiong, Jinjun
Formato: Preprint
Publicado: 2025
Materias:
Acceso en línea:
Etiquetas: Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
_version_ 1866914522494664704
author Xu, Chenhui
Yu, Fuxun
Bianco, Michael J.
Kovarskiy, Jacob
Tang, Raphael
Zhang, Qi
Xu, Zirui
LeVine, Will
Dubbs, Brandon
Liao, Heming
Burgess, Cassandra
Bag, Suvam
Patravali, Jay
Kukal, Rupanjali
Figueroa, Mikael
Madhok, Rishi
Karianakis, Nikolaos
Xiong, Jinjun
author_facet Xu, Chenhui
Yu, Fuxun
Bianco, Michael J.
Kovarskiy, Jacob
Tang, Raphael
Zhang, Qi
Xu, Zirui
LeVine, Will
Dubbs, Brandon
Liao, Heming
Burgess, Cassandra
Bag, Suvam
Patravali, Jay
Kukal, Rupanjali
Figueroa, Mikael
Madhok, Rishi
Karianakis, Nikolaos
Xiong, Jinjun
contents Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. While raw geospatial imagery is abundant, the amount of task-direct supervision falls far behind that of common domains. In this work, we validate an important conclusion: indirect verifiable rewards, derived from seemingly unrelated metadata, are sufficient to induce sophisticated and generalizable geospatial reasoning across a wide range of downstream tasks (25+). We present Geo-R1 as one empirical instantiation of this paradigm. Rather than relying on limited task-specific annotations (i.e., direct rewards), Geo-R1 utilizes scalable, verifiable indirect proxy rewards based on cross-view alignment with metadata (geolocation information) to drive reinforcement learning at scale. Such indirect rewards successfully motivate the model to discover and internalize zero-shot geospatial reasoning across diverse tasks, achieving extraordinary zero-shot transfer on out-of-distribution benchmarks and even surpassing fully supervised specialists on certain benchmarks. These findings indicate that optimizing for indirect verifiable rewards may provide a scalable pathway to unlock generalized reasoning capabilities in rare domains with massive unlabeled data archives. Our code is availavle at: https://github.com/miniHuiHui/Geo-R1.
format Preprint
id arxiv_https___arxiv_org_abs_2510_00072
institution arXiv
publishDate 2025
record_format arxiv
spellingShingle Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
Xu, Chenhui
Yu, Fuxun
Bianco, Michael J.
Kovarskiy, Jacob
Tang, Raphael
Zhang, Qi
Xu, Zirui
LeVine, Will
Dubbs, Brandon
Liao, Heming
Burgess, Cassandra
Bag, Suvam
Patravali, Jay
Kukal, Rupanjali
Figueroa, Mikael
Madhok, Rishi
Karianakis, Nikolaos
Xiong, Jinjun
Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
Training robust reasoning vision-language models (VLMs) in rare domains (such as geospatial) is fundamentally constrained by supervision scarcity. While raw geospatial imagery is abundant, the amount of task-direct supervision falls far behind that of common domains. In this work, we validate an important conclusion: indirect verifiable rewards, derived from seemingly unrelated metadata, are sufficient to induce sophisticated and generalizable geospatial reasoning across a wide range of downstream tasks (25+). We present Geo-R1 as one empirical instantiation of this paradigm. Rather than relying on limited task-specific annotations (i.e., direct rewards), Geo-R1 utilizes scalable, verifiable indirect proxy rewards based on cross-view alignment with metadata (geolocation information) to drive reinforcement learning at scale. Such indirect rewards successfully motivate the model to discover and internalize zero-shot geospatial reasoning across diverse tasks, achieving extraordinary zero-shot transfer on out-of-distribution benchmarks and even surpassing fully supervised specialists on certain benchmarks. These findings indicate that optimizing for indirect verifiable rewards may provide a scalable pathway to unlock generalized reasoning capabilities in rare domains with massive unlabeled data archives. Our code is availavle at: https://github.com/miniHuiHui/Geo-R1.
title Unlocking Zero-Shot Geospatial Reasoning via Indirect Rewards
topic Computer Vision and Pattern Recognition
Artificial Intelligence
Machine Learning
url https://arxiv.org/abs/2510.00072