RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
Fuente:
arXiv
Saved in:
| Main Authors: | Zhuang, Qiyuan, Xu, He-Yang, Wang, Yijun, Zhao, Xin-Yang, Li, Yang-Yang, Wei, Xiu-Shen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
General Flow as Foundation Affordance for Scalable Robot Learning
by: Yuan, Chengbo, et al.
Published: (2024)
by: Yuan, Chengbo, et al.
Published: (2024)
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
by: Maksutova, Aiza, et al.
Published: (2026)
by: Maksutova, Aiza, et al.
Published: (2026)
Cross-Modal Visual Relocalization in Prior LiDAR Maps Utilizing Intensity Textures
by: Shen, Qiyuan, et al.
Published: (2024)
by: Shen, Qiyuan, et al.
Published: (2024)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
by: Qi, Xiuxiu, et al.
Published: (2025)
by: Qi, Xiuxiu, et al.
Published: (2025)
ClothPPO: A Proximal Policy Optimization Enhancing Framework for Robotic Cloth Manipulation with Observation-Aligned Action Spaces
by: Yang, Libing, et al.
Published: (2024)
by: Yang, Libing, et al.
Published: (2024)
BEACON: Language-Conditioned Navigation Affordance Prediction under Occlusion
by: Gao, Xinyu, et al.
Published: (2026)
by: Gao, Xinyu, et al.
Published: (2026)
VLA-Pro: Cross-Task Procedural Memory Transfer for Vision-Language-Action Models
by: Si, Shengyu, et al.
Published: (2026)
by: Si, Shengyu, et al.
Published: (2026)
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
GarmentPile: Point-Level Visual Affordance Guided Retrieval and Adaptation for Cluttered Garments Manipulation
by: Wu, Ruihai, et al.
Published: (2025)
by: Wu, Ruihai, et al.
Published: (2025)
A Light-Weight Framework for Open-Set Object Detection with Decoupled Feature Alignment in Joint Space
by: He, Yonghao, et al.
Published: (2024)
by: He, Yonghao, et al.
Published: (2024)
RoboPoint: A Vision-Language Model for Spatial Affordance Prediction for Robotics
by: Yuan, Wentao, et al.
Published: (2024)
by: Yuan, Wentao, et al.
Published: (2024)
Grounding 3D Object Affordance with Language Instructions, Visual Observations and Interactions
by: Zhu, He, et al.
Published: (2025)
by: Zhu, He, et al.
Published: (2025)
DVGT-2: Vision-Geometry-Action Model for Autonomous Driving at Scale
by: Zuo, Sicheng, et al.
Published: (2026)
by: Zuo, Sicheng, et al.
Published: (2026)
LoopVLA: Learning Sufficiency in Recurrent Refinement for Vision-Language-Action Models
by: Shen, Boyang, et al.
Published: (2026)
by: Shen, Boyang, et al.
Published: (2026)
ROSA: Harnessing Robot States for Vision-Language and Action Alignment
by: Wen, Yuqing, et al.
Published: (2025)
by: Wen, Yuqing, et al.
Published: (2025)
Self-Correcting VLA: Online Action Refinement via Sparse World Imagination
by: Liu, Chenyv, et al.
Published: (2026)
by: Liu, Chenyv, et al.
Published: (2026)
Learning Environment-Aware Affordance for 3D Articulated Object Manipulation under Occlusions
by: Wu, Ruihai, et al.
Published: (2023)
by: Wu, Ruihai, et al.
Published: (2023)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
TRACER: Texture-Robust Affordance Chain-of-Thought for Deformable-Object Refinement
by: Jia, Wanjun, et al.
Published: (2026)
by: Jia, Wanjun, et al.
Published: (2026)
LoD-Loc v3: Generalized Aerial Localization in Dense Cities using Instance Silhouette Alignment
by: Peng, Shuaibang, et al.
Published: (2026)
by: Peng, Shuaibang, et al.
Published: (2026)
A Survey on Vision-Language-Action Models for Autonomous Driving
by: Jiang, Sicong, et al.
Published: (2025)
by: Jiang, Sicong, et al.
Published: (2025)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
by: Wu, Xiaofei, et al.
Published: (2026)
by: Wu, Xiaofei, et al.
Published: (2026)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
by: Kuang, Yuxuan, et al.
Published: (2024)
by: Kuang, Yuxuan, et al.
Published: (2024)
Visual Affordance Prediction: Survey and Reproducibility
by: Apicella, Tommaso, et al.
Published: (2025)
by: Apicella, Tommaso, et al.
Published: (2025)
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
by: Tang, Yihe, et al.
Published: (2025)
by: Tang, Yihe, et al.
Published: (2025)
KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System
by: Xia, Zhongyu, et al.
Published: (2025)
by: Xia, Zhongyu, et al.
Published: (2025)
ALAM: Algebraically Consistent Latent Action Model for Vision-Language-Action Models
by: Tang, Zuojin, et al.
Published: (2026)
by: Tang, Zuojin, et al.
Published: (2026)
Efficient Alignment of Unconditioned Action Prior for Language-conditioned Pick and Place in Clutter
by: Xu, Kechun, et al.
Published: (2025)
by: Xu, Kechun, et al.
Published: (2025)
EA-WM: Event-Aware Generative World Model with Structured Kinematic-to-Visual Action Fields
by: Yang, Zhaoyang, et al.
Published: (2026)
by: Yang, Zhaoyang, et al.
Published: (2026)
Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
by: Morin, Sacha, et al.
Published: (2025)
by: Morin, Sacha, et al.
Published: (2025)
From Pixels to Tokens: A Systematic Study of Latent Action Supervision for Vision-Language-Action Models
by: Lin, Yihan, et al.
Published: (2026)
by: Lin, Yihan, et al.
Published: (2026)
Benefit from Reference: Retrieval-Augmented Cross-modal Point Cloud Completion
by: Hou, Hongye, et al.
Published: (2025)
by: Hou, Hongye, et al.
Published: (2025)
SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
by: Wu, Shengkai, et al.
Published: (2025)
by: Wu, Shengkai, et al.
Published: (2025)
Attention-Enhanced Cross-modal Localization Between 360 Images and Point Clouds
by: Zhao, Zhipeng, et al.
Published: (2022)
by: Zhao, Zhipeng, et al.
Published: (2022)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments
by: Zhu, Guoliang, et al.
Published: (2026)
by: Zhu, Guoliang, et al.
Published: (2026)
MemoNav: Working Memory Model for Visual Navigation
by: Li, Hongxin, et al.
Published: (2024)
by: Li, Hongxin, et al.
Published: (2024)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting
by: Lin, Juyi, et al.
Published: (2025)
by: Lin, Juyi, et al.
Published: (2025)
RAP: Retrieval-Augmented Planner for Adaptive Procedure Planning in Instructional Videos
by: Zare, Ali, et al.
Published: (2024)
by: Zare, Ali, et al.
Published: (2024)
Similar Items
-
General Flow as Foundation Affordance for Scalable Robot Learning
by: Yuan, Chengbo, et al.
Published: (2024) -
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
by: Maksutova, Aiza, et al.
Published: (2026) -
Cross-Modal Visual Relocalization in Prior LiDAR Maps Utilizing Intensity Textures
by: Shen, Qiyuan, et al.
Published: (2024) -
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
by: Qi, Xiuxiu, et al.
Published: (2025) -
ClothPPO: A Proximal Policy Optimization Enhancing Framework for Robotic Cloth Manipulation with Observation-Aligned Action Spaces
by: Yang, Libing, et al.
Published: (2024)