GaussFly: Contrastive Reinforcement Learning for Visuomotor Policies in 3D Gaussian Fields

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Zhang, Yuhang, Li, Mingsheng, Shang, Yujing, Yu, Zhuoyuan, Yan, Chao, Xiao, Jiaping, Feroskhan, Mir
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866913009283104768
author Zhang, Yuhang
Li, Mingsheng
Shang, Yujing
Yu, Zhuoyuan
Yan, Chao
Xiao, Jiaping
Feroskhan, Mir
author_facet Zhang, Yuhang
Li, Mingsheng
Shang, Yujing
Yu, Zhuoyuan
Yan, Chao
Xiao, Jiaping
Feroskhan, Mir
contents Learning visuomotor policies for Autonomous Aerial Vehicles (AAVs) relying solely on monocular vision is an attractive yet highly challenging paradigm. Existing end-to-end learning approaches directly map high-dimensional RGB observations to action commands, which frequently suffer from low sample efficiency and severe sim-to-real gaps due to the visual discrepancy between simulation and physical domains. To address these long-standing challenges, we propose GaussFly, a novel framework that explicitly decouples representation learning from policy optimization through a cohesive real-to-sim-to-real paradigm. First, to achieve a high-fidelity real-to-sim transition, we reconstruct training scenes using 3D Gaussian Splatting (3DGS) augmented with explicit geometric constraints. Second, to ensure robust sim-to-real transfer, we leverage these photorealistic simulated environments and employ contrastive representation learning to extract compact, noise-resilient latent features from the rendered RGB images. By utilizing this pre-trained encoder to provide low-dimensional feature inputs, the computational burden on the visuomotor policy is significantly reduced while its resistance against visual noise is inherently enhanced. Extensive experiments in simulated and real-world environments demonstrate that GaussFly achieves superior sample efficiency and asymptotic performance compared to baselines. Crucially, it enables robust and zero-shot policy transfer to unseen real-world environments with complex textures, effectively bridging the sim-to-real gap.
format Preprint
id arxiv_https___arxiv_org_abs_2604_05062
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GaussFly: Contrastive Reinforcement Learning for Visuomotor Policies in 3D Gaussian Fields
Zhang, Yuhang
Li, Mingsheng
Shang, Yujing
Yu, Zhuoyuan
Yan, Chao
Xiao, Jiaping
Feroskhan, Mir
Robotics
Learning visuomotor policies for Autonomous Aerial Vehicles (AAVs) relying solely on monocular vision is an attractive yet highly challenging paradigm. Existing end-to-end learning approaches directly map high-dimensional RGB observations to action commands, which frequently suffer from low sample efficiency and severe sim-to-real gaps due to the visual discrepancy between simulation and physical domains. To address these long-standing challenges, we propose GaussFly, a novel framework that explicitly decouples representation learning from policy optimization through a cohesive real-to-sim-to-real paradigm. First, to achieve a high-fidelity real-to-sim transition, we reconstruct training scenes using 3D Gaussian Splatting (3DGS) augmented with explicit geometric constraints. Second, to ensure robust sim-to-real transfer, we leverage these photorealistic simulated environments and employ contrastive representation learning to extract compact, noise-resilient latent features from the rendered RGB images. By utilizing this pre-trained encoder to provide low-dimensional feature inputs, the computational burden on the visuomotor policy is significantly reduced while its resistance against visual noise is inherently enhanced. Extensive experiments in simulated and real-world environments demonstrate that GaussFly achieves superior sample efficiency and asymptotic performance compared to baselines. Crucially, it enables robust and zero-shot policy transfer to unseen real-world environments with complex textures, effectively bridging the sim-to-real gap.
title GaussFly: Contrastive Reinforcement Learning for Visuomotor Policies in 3D Gaussian Fields
topic Robotics
url https://arxiv.org/abs/2604.05062