ParkourFormer: Integrating Predictive Supervision and Sequence Modeling into Parkour Locomotion

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Mai, Yanheng, Xu, Wenhao, Huang, Zirui, Fu, Yifei, Dong, Shengwei, Wang, Xinjue, Huang, Kailun, Xie, Yanzhe, Xu, Renjing
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866914601848799232
author Mai, Yanheng
Xu, Wenhao
Huang, Zirui
Fu, Yifei
Dong, Shengwei
Wang, Xinjue
Huang, Kailun
Xie, Yanzhe
Xu, Renjing
author_facet Mai, Yanheng
Xu, Wenhao
Huang, Zirui
Fu, Yifei
Dong, Shengwei
Wang, Xinjue
Huang, Kailun
Xie, Yanzhe
Xu, Renjing
contents Humanoid parkour requires locomotion policies to coordinate whole-body dynamics across rapidly changing terrains such as stairs, gaps, slopes, and obstacles. Existing reinforcement learning policies are largely reactive, mapping observations directly to actions without explicitly modeling future body states. Such modeling becomes critical in agile locomotion tasks where successful motion execution depends strongly on anticipating upcoming contact transitions and body dynamics. We present ParkourFormer, a Transformer-based sequence modeling framework that reformulates humanoid locomotion as a future-conditioned decision-making problem. The current robot state queries historical sensorimotor trajectories through cross-attention, while a lightweight prediction head forecasts short-horizon future proprioceptive states. The predicted future states, trained with supervised signals, are fused with temporal features to generate actions, enabling the policy to jointly reason over motion history and anticipated future dynamics. We evaluate ParkourFormer on a diverse multi-terrain humanoid parkour benchmark including stairs, gaps, slopes, rough terrain, and obstacle traversal. Experiments in simulation and on a real humanoid robot show that ParkourFormer achieves a 93.85% average traversal success rate on highly challenging terrains, with improvements of up to 42.73% over strong MLP, MoE-based MLP, and vanilla Transformer baselines, while maintaining a single unified policy across all terrain types. These results demonstrate that explicit future-state modeling significantly improves robustness and generalization for agile whole-body locomotion.
format Preprint
id arxiv_https___arxiv_org_abs_2605_25782
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle ParkourFormer: Integrating Predictive Supervision and Sequence Modeling into Parkour Locomotion
Mai, Yanheng
Xu, Wenhao
Huang, Zirui
Fu, Yifei
Dong, Shengwei
Wang, Xinjue
Huang, Kailun
Xie, Yanzhe
Xu, Renjing
Robotics
Humanoid parkour requires locomotion policies to coordinate whole-body dynamics across rapidly changing terrains such as stairs, gaps, slopes, and obstacles. Existing reinforcement learning policies are largely reactive, mapping observations directly to actions without explicitly modeling future body states. Such modeling becomes critical in agile locomotion tasks where successful motion execution depends strongly on anticipating upcoming contact transitions and body dynamics. We present ParkourFormer, a Transformer-based sequence modeling framework that reformulates humanoid locomotion as a future-conditioned decision-making problem. The current robot state queries historical sensorimotor trajectories through cross-attention, while a lightweight prediction head forecasts short-horizon future proprioceptive states. The predicted future states, trained with supervised signals, are fused with temporal features to generate actions, enabling the policy to jointly reason over motion history and anticipated future dynamics. We evaluate ParkourFormer on a diverse multi-terrain humanoid parkour benchmark including stairs, gaps, slopes, rough terrain, and obstacle traversal. Experiments in simulation and on a real humanoid robot show that ParkourFormer achieves a 93.85% average traversal success rate on highly challenging terrains, with improvements of up to 42.73% over strong MLP, MoE-based MLP, and vanilla Transformer baselines, while maintaining a single unified policy across all terrain types. These results demonstrate that explicit future-state modeling significantly improves robustness and generalization for agile whole-body locomotion.
title ParkourFormer: Integrating Predictive Supervision and Sequence Modeling into Parkour Locomotion
topic Robotics
url https://arxiv.org/abs/2605.25782