VisFly-Lab: Unified Differentiable Framework for First-Order Reinforcement Learning of Quadrotor Control

Fuente: arXiv
Gespeichert in:
Bibliographische Detailangaben
Hauptverfasser: Li, Fanxing, Sun, Fangyu, Zhang, Tianbao, Wu, Shuyu, Zuo, Dexin, Yan, yufei, Yu, Wenxian, Zou, Danping
Format: Preprint
Veröffentlicht: 2026
Schlagworte:
Online-Zugang:
Tags: Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
_version_ 1866908905630597120
author Li, Fanxing
Sun, Fangyu
Zhang, Tianbao
Wu, Shuyu
Zuo, Dexin
Yan, yufei
Yu, Wenxian
Zou, Danping
author_facet Li, Fanxing
Sun, Fangyu
Zhang, Tianbao
Wu, Shuyu
Zuo, Dexin
Yan, yufei
Yu, Wenxian
Zou, Danping
contents First-order reinforcement learning with differentiable simulation is promising for quadrotor control, but practical progress remains fragmented across task-specific settings. To support more systematic development and evaluation, we present a unified differentiable framework for multi-task quadrotor control. The framework is wrapped, extensible, and equipped with deployment-oriented dynamics, providing a common interface across four representative tasks: hovering, tracking, landing, and racing. We also present the suite of first-order learning algorithms, where we identify two practical bottlenecks of standard first-order training: limited state coverage caused by horizon initialization and gradient bias caused by partially non-differentiable rewards. To address these issues, we propose Amended Backpropagation Through Time (ABPT), which combines differentiable rollout optimization, a value-based auxiliary objective, and visited-state initialization to improve training robustness. Experimental results show that ABPT yields the clearest gains in tasks with partially non-differentiable rewards, while remaining competitive in fully differentiable settings. We further provide proof-of-concept real-world deployments showing initial transferability of policies learned in the proposed framework beyond simulation.
format Preprint
id arxiv_https___arxiv_org_abs_2603_21123
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle VisFly-Lab: Unified Differentiable Framework for First-Order Reinforcement Learning of Quadrotor Control
Li, Fanxing
Sun, Fangyu
Zhang, Tianbao
Wu, Shuyu
Zuo, Dexin
Yan, yufei
Yu, Wenxian
Zou, Danping
Robotics
First-order reinforcement learning with differentiable simulation is promising for quadrotor control, but practical progress remains fragmented across task-specific settings. To support more systematic development and evaluation, we present a unified differentiable framework for multi-task quadrotor control. The framework is wrapped, extensible, and equipped with deployment-oriented dynamics, providing a common interface across four representative tasks: hovering, tracking, landing, and racing. We also present the suite of first-order learning algorithms, where we identify two practical bottlenecks of standard first-order training: limited state coverage caused by horizon initialization and gradient bias caused by partially non-differentiable rewards. To address these issues, we propose Amended Backpropagation Through Time (ABPT), which combines differentiable rollout optimization, a value-based auxiliary objective, and visited-state initialization to improve training robustness. Experimental results show that ABPT yields the clearest gains in tasks with partially non-differentiable rewards, while remaining competitive in fully differentiable settings. We further provide proof-of-concept real-world deployments showing initial transferability of policies learned in the proposed framework beyond simulation.
title VisFly-Lab: Unified Differentiable Framework for First-Order Reinforcement Learning of Quadrotor Control
topic Robotics
url https://arxiv.org/abs/2603.21123