$π$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Tucker, Johnathan, Liu, Denis, Swann, Aiden, Ren, Allen, Yu, Javier, Sun, Jiankai, Kim, Brandon, McGranahan, Lachlain, Vuong, Quan, Schwager, Mac
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866918409668657152
author Tucker, Johnathan
Liu, Denis
Swann, Aiden
Ren, Allen
Yu, Javier
Sun, Jiankai
Kim, Brandon
McGranahan, Lachlain
Vuong, Quan
Schwager, Mac
author_facet Tucker, Johnathan
Liu, Denis
Swann, Aiden
Ren, Allen
Yu, Javier
Sun, Jiankai
Kim, Brandon
McGranahan, Lachlain
Vuong, Quan
Schwager, Mac
contents Vision-Language-Action (VLA) models such as $π_0$ have demonstrated remarkable generalization across diverse fixed-base manipulators. However, transferring these foundation models to aerial platforms remains an open challenge due to the fundamental mismatch between the quasi-static dynamics of fixed-base arms and the underactuated, highly dynamic nature of flight. In this work, we introduce AirVLA, a system that investigates the transferability of manipulation-pretrained VLAs to aerial pick-and-place tasks. We find that while visual representations transfer effectively, the specific control dynamics required for flight do not. To bridge this "dynamics gap" without retraining the foundation model, we introduce a Payload-Aware Guidance mechanism that injects payload constraints directly into the policy's flow-matching sampling process. To overcome data scarcity, we further utilize a Gaussian Splatting pipeline to synthesize navigation training data. We evaluate our method through a cumulative 460 real-world experiments which demonstrate that this synthetic data is a key enabler of performance, unlocking 100% success in navigation tasks where directly fine-tuning on teleoperation data alone attains 81% success. Our inference-time intervention, Payload-Aware Guidance, increases real-world pick-and-place task success from 23% to 50%. Finally, we evaluate the model on a long-horizon compositional task, achieving a 62% overall success rate. These results suggest that pre-trained manipulation VLAs, with appropriate data augmentation and physics-informed guidance, can transfer to aerial manipulation and navigation, as well as the composition of these tasks.
format Preprint
id arxiv_https___arxiv_org_abs_2603_25038
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle $π$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial Manipulation
Tucker, Johnathan
Liu, Denis
Swann, Aiden
Ren, Allen
Yu, Javier
Sun, Jiankai
Kim, Brandon
McGranahan, Lachlain
Vuong, Quan
Schwager, Mac
Robotics
Vision-Language-Action (VLA) models such as $π_0$ have demonstrated remarkable generalization across diverse fixed-base manipulators. However, transferring these foundation models to aerial platforms remains an open challenge due to the fundamental mismatch between the quasi-static dynamics of fixed-base arms and the underactuated, highly dynamic nature of flight. In this work, we introduce AirVLA, a system that investigates the transferability of manipulation-pretrained VLAs to aerial pick-and-place tasks. We find that while visual representations transfer effectively, the specific control dynamics required for flight do not. To bridge this "dynamics gap" without retraining the foundation model, we introduce a Payload-Aware Guidance mechanism that injects payload constraints directly into the policy's flow-matching sampling process. To overcome data scarcity, we further utilize a Gaussian Splatting pipeline to synthesize navigation training data. We evaluate our method through a cumulative 460 real-world experiments which demonstrate that this synthetic data is a key enabler of performance, unlocking 100% success in navigation tasks where directly fine-tuning on teleoperation data alone attains 81% success. Our inference-time intervention, Payload-Aware Guidance, increases real-world pick-and-place task success from 23% to 50%. Finally, we evaluate the model on a long-horizon compositional task, achieving a 62% overall success rate. These results suggest that pre-trained manipulation VLAs, with appropriate data augmentation and physics-informed guidance, can transfer to aerial manipulation and navigation, as well as the composition of these tasks.
title $π$, But Make It Fly: Physics-Guided Transfer of VLA Models to Aerial Manipulation
topic Robotics
url https://arxiv.org/abs/2603.25038