PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Liu, Yuanzhe, Zhu, Jingyuan, Mo, Yuchen, Li, Gen, Cao, Xu, Jin, Jin, Shen, Yifan, Li, Zhengyuan, Yu, Tianjiao, Yuan, Wenzhen, Ding, Fangqiang, Lourentzou, Ismini
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866915914577870848
author Liu, Yuanzhe
Zhu, Jingyuan
Mo, Yuchen
Li, Gen
Cao, Xu
Jin, Jin
Shen, Yifan
Li, Zhengyuan
Yu, Tianjiao
Yuan, Wenzhen
Ding, Fangqiang
Lourentzou, Ismini
author_facet Liu, Yuanzhe
Zhu, Jingyuan
Mo, Yuchen
Li, Gen
Cao, Xu
Jin, Jin
Shen, Yifan
Li, Zhengyuan
Yu, Tianjiao
Yuan, Wenzhen
Ding, Fangqiang
Lourentzou, Ismini
contents Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a subtask, leading to critical execution errors such as repeated actions, missed steps, and premature termination. To address these challenges, we introduce PALM, a VLA framework that structures policy learning around interaction-centric affordance reasoning and subtask progress cues. PALM distills complementary affordance representations that capture object relevance, contact geometry, spatial placements, and motion dynamics, and serve as task-relevant anchors for visuomotor control. To further stabilize long-horizon execution, PALM predicts continuous within-subtask progress, enabling seamless subtask transitions. Across extensive simulation and real-world experiments, PALM consistently outperforms baselines, achieving a 91.8% success rate on LIBERO-LONG, a 12.5% improvement in average length on CALVIN ABC->D, and a 2x improvement over real-world baselines across three long-horizon generalization settings.
format Preprint
id arxiv_https___arxiv_org_abs_2601_07060
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
Liu, Yuanzhe
Zhu, Jingyuan
Mo, Yuchen
Li, Gen
Cao, Xu
Jin, Jin
Shen, Yifan
Li, Zhengyuan
Yu, Tianjiao
Yuan, Wenzhen
Ding, Fangqiang
Lourentzou, Ismini
Robotics
Recent advancements in vision-language-action (VLA) models have shown promise in robotic manipulation, yet they continue to struggle with long-horizon, multi-step tasks. Existing methods lack internal reasoning mechanisms that can identify task-relevant interaction cues or track progress within a subtask, leading to critical execution errors such as repeated actions, missed steps, and premature termination. To address these challenges, we introduce PALM, a VLA framework that structures policy learning around interaction-centric affordance reasoning and subtask progress cues. PALM distills complementary affordance representations that capture object relevance, contact geometry, spatial placements, and motion dynamics, and serve as task-relevant anchors for visuomotor control. To further stabilize long-horizon execution, PALM predicts continuous within-subtask progress, enabling seamless subtask transitions. Across extensive simulation and real-world experiments, PALM consistently outperforms baselines, achieving a 91.8% success rate on LIBERO-LONG, a 12.5% improvement in average length on CALVIN ABC->D, and a 2x improvement over real-world baselines across three long-horizon generalization settings.
title PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
topic Robotics
url https://arxiv.org/abs/2601.07060