PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Yuanzhe, Zhu, Jingyuan, Mo, Yuchen, Li, Gen, Cao, Xu, Jin, Jin, Shen, Yifan, Li, Zhengyuan, Yu, Tianjiao, Yuan, Wenzhen, Ding, Fangqiang, Lourentzou, Ismini |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
by: Zhou, Xiaona, et al.
Published: (2026)
by: Zhou, Xiaona, et al.
Published: (2026)
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025)
by: Shen, Yifan, et al.
Published: (2025)
Commonsense for Zero-Shot Natural Language Video Localization
by: Holla, Meghana, et al.
Published: (2023)
by: Holla, Meghana, et al.
Published: (2023)
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
by: Yu, Tianjiao, et al.
Published: (2026)
by: Yu, Tianjiao, et al.
Published: (2026)
DoorBot: Closed-Loop Task Planning and Manipulation for Door Opening in the Wild with Haptic Feedback
by: Wang, Zhi, et al.
Published: (2025)
by: Wang, Zhi, et al.
Published: (2025)
Hierarchical Dataset Selection for High-Quality Data Sharing
by: Zhou, Xiaona, et al.
Published: (2025)
by: Zhou, Xiaona, et al.
Published: (2025)
FAIR: Facilitating Artificial Intelligence Resilience in Manufacturing Industrial Internet
by: Zeng, Yingyan, et al.
Published: (2025)
by: Zeng, Yingyan, et al.
Published: (2025)
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
by: Li, Xinzhuo, et al.
Published: (2025)
by: Li, Xinzhuo, et al.
Published: (2025)
CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
by: Nguyen, Kiet A., et al.
Published: (2024)
by: Nguyen, Kiet A., et al.
Published: (2024)
Learning to Double Guess: An Active Perception Approach for Estimating the Center of Mass of Arbitrary Objects
by: Jin, Shengmiao, et al.
Published: (2025)
by: Jin, Shengmiao, et al.
Published: (2025)
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation
by: Hu, Yuxuan, et al.
Published: (2026)
by: Hu, Yuxuan, et al.
Published: (2026)
Toward Cognitive Supersensing in Multimodal Large Language Model
by: Li, Boyi, et al.
Published: (2026)
by: Li, Boyi, et al.
Published: (2026)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
by: Ogunleye, Makanjuola, et al.
Published: (2026)
by: Ogunleye, Makanjuola, et al.
Published: (2026)
mTSBench: Benchmarking Multivariate Time Series Anomaly Detection and Model Selection at Scale
by: Zhou, Xiaona, et al.
Published: (2025)
by: Zhou, Xiaona, et al.
Published: (2025)
EgoForge: Goal-Directed Egocentric World Simulator
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
Part$^{2}$GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
Thinking in Text and Images: Interleaved Vision--Language Reasoning Traces for Long-Horizon Robot Manipulation
by: Liu, Jinkun, et al.
Published: (2026)
by: Liu, Jinkun, et al.
Published: (2026)
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
by: Li, Gen, et al.
Published: (2024)
by: Li, Gen, et al.
Published: (2024)
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
by: Wahed, Muntasir, et al.
Published: (2024)
by: Wahed, Muntasir, et al.
Published: (2024)
MMPlanner: Zero-Shot Multimodal Procedural Planning with Chain-of-Thought Object State Reasoning
by: Tabassum, Afrina, et al.
Published: (2025)
by: Tabassum, Afrina, et al.
Published: (2025)
TRACE: Textual Reasoning for Affordance Coordinate Extraction
by: Park, Sangyun, et al.
Published: (2025)
by: Park, Sangyun, et al.
Published: (2025)
LongBench: Evaluating Robotic Manipulation Policies on Real-World Long-Horizon Tasks
by: Chen, Xueyao, et al.
Published: (2026)
by: Chen, Xueyao, et al.
Published: (2026)
Sensor-Invariant Tactile Representation
by: Gupta, Harsh, et al.
Published: (2025)
by: Gupta, Harsh, et al.
Published: (2025)
RT-Affordance: Affordances are Versatile Intermediate Representations for Robot Manipulation
by: Nasiriany, Soroush, et al.
Published: (2024)
by: Nasiriany, Soroush, et al.
Published: (2024)
SAGE: Scene Graph-Aware Guidance and Execution for Long-Horizon Manipulation Tasks
by: Li, Jialiang, et al.
Published: (2025)
by: Li, Jialiang, et al.
Published: (2025)
RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance
by: Venkatesh, Kavana, et al.
Published: (2024)
by: Venkatesh, Kavana, et al.
Published: (2024)
ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
by: Shen, Ying, et al.
Published: (2023)
by: Shen, Ying, et al.
Published: (2023)
Uncertainty in Action: Confidence Elicitation in Embodied Agents
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
A0: An Affordance-Aware Hierarchical Model for General Robotic Manipulation
by: Xu, Rongtao, et al.
Published: (2025)
by: Xu, Rongtao, et al.
Published: (2025)
SARM: Stage-Aware Reward Modeling for Long Horizon Robot Manipulation
by: Chen, Qianzhong, et al.
Published: (2025)
by: Chen, Qianzhong, et al.
Published: (2025)
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
by: Tang, Yihe, et al.
Published: (2025)
by: Tang, Yihe, et al.
Published: (2025)
Non-Markovian Long-Horizon Robot Manipulation via Keyframe Chaining
by: Chen, Yipeng, et al.
Published: (2026)
by: Chen, Yipeng, et al.
Published: (2026)
uaMix-MAE: Efficient Tuning of Pretrained Audio Transformers with Unsupervised Audio Mixtures
by: Tabassum, Afrina, et al.
Published: (2024)
by: Tabassum, Afrina, et al.
Published: (2024)
AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
by: Li, Mingyang, et al.
Published: (2026)
by: Li, Mingyang, et al.
Published: (2026)
AnchorDP3: 3D Affordance Guided Sparse Diffusion Policy for Robotic Manipulation
by: Zhao, Ziyan, et al.
Published: (2025)
by: Zhao, Ziyan, et al.
Published: (2025)
Evaluating Cognitive Age Alignment in Interactive AI Agents
by: Shen, Yifan, et al.
Published: (2026)
by: Shen, Yifan, et al.
Published: (2026)
ProgressVLA: Progress-Guided Diffusion Policy for Vision-Language Robotic Manipulation
by: Yan, Hongyu, et al.
Published: (2026)
by: Yan, Hongyu, et al.
Published: (2026)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
by: Wahed, Muntasir, et al.
Published: (2025)
by: Wahed, Muntasir, et al.
Published: (2025)
Similar Items
-
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
by: Yu, Tianjiao, et al.
Published: (2025) -
Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
by: Zhou, Xiaona, et al.
Published: (2026) -
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
by: Shen, Ying, et al.
Published: (2026) -
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs
by: Shen, Yifan, et al.
Published: (2025) -
Commonsense for Zero-Shot Natural Language Video Localization
by: Holla, Meghana, et al.
Published: (2023)