VisFly-Lab: Unified Differentiable Framework for First-Order Reinforcement Learning of Quadrotor Control
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | , , , , , , , |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
| _version_ | 1866908905630597120 |
|---|---|
| author | Li, Fanxing Sun, Fangyu Zhang, Tianbao Wu, Shuyu Zuo, Dexin Yan, yufei Yu, Wenxian Zou, Danping |
| author_facet | Li, Fanxing Sun, Fangyu Zhang, Tianbao Wu, Shuyu Zuo, Dexin Yan, yufei Yu, Wenxian Zou, Danping |
| contents | First-order reinforcement learning with differentiable simulation is promising for quadrotor control, but practical progress remains fragmented across task-specific settings. To support more systematic development and evaluation, we present a unified differentiable framework for multi-task quadrotor control. The framework is wrapped, extensible, and equipped with deployment-oriented dynamics, providing a common interface across four representative tasks: hovering, tracking, landing, and racing. We also present the suite of first-order learning algorithms, where we identify two practical bottlenecks of standard first-order training: limited state coverage caused by horizon initialization and gradient bias caused by partially non-differentiable rewards. To address these issues, we propose Amended Backpropagation Through Time (ABPT), which combines differentiable rollout optimization, a value-based auxiliary objective, and visited-state initialization to improve training robustness. Experimental results show that ABPT yields the clearest gains in tasks with partially non-differentiable rewards, while remaining competitive in fully differentiable settings. We further provide proof-of-concept real-world deployments showing initial transferability of policies learned in the proposed framework beyond simulation. |
| format | Preprint |
| id |
arxiv_https___arxiv_org_abs_2603_21123 |
| institution | arXiv |
| publishDate | 2026 |
| record_format | arxiv |
| spellingShingle | VisFly-Lab: Unified Differentiable Framework for First-Order Reinforcement Learning of Quadrotor Control Li, Fanxing Sun, Fangyu Zhang, Tianbao Wu, Shuyu Zuo, Dexin Yan, yufei Yu, Wenxian Zou, Danping Robotics First-order reinforcement learning with differentiable simulation is promising for quadrotor control, but practical progress remains fragmented across task-specific settings. To support more systematic development and evaluation, we present a unified differentiable framework for multi-task quadrotor control. The framework is wrapped, extensible, and equipped with deployment-oriented dynamics, providing a common interface across four representative tasks: hovering, tracking, landing, and racing. We also present the suite of first-order learning algorithms, where we identify two practical bottlenecks of standard first-order training: limited state coverage caused by horizon initialization and gradient bias caused by partially non-differentiable rewards. To address these issues, we propose Amended Backpropagation Through Time (ABPT), which combines differentiable rollout optimization, a value-based auxiliary objective, and visited-state initialization to improve training robustness. Experimental results show that ABPT yields the clearest gains in tasks with partially non-differentiable rewards, while remaining competitive in fully differentiable settings. We further provide proof-of-concept real-world deployments showing initial transferability of policies learned in the proposed framework beyond simulation. |
| title | VisFly-Lab: Unified Differentiable Framework for First-Order Reinforcement Learning of Quadrotor Control |
| topic | Robotics |
| url | https://arxiv.org/abs/2603.21123 |