Afford-VLA: Action-Aligned Visual Planning via Internalized Affordance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Runze, Fu, Yuqian, Li, Yu, Lin, Tao, Qian, Tianwen, Elhoseiny, Mohamed, Zhao, Bo, Fu, Yanwei, Jiang, Yu-Gang, Xue, Xiangyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
von: Kong, Weijie, et al.
Veröffentlicht: (2026)
von: Kong, Weijie, et al.
Veröffentlicht: (2026)
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
von: Liu, Zhenyang, et al.
Veröffentlicht: (2026)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2026)
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
von: Wang, Kuanning, et al.
Veröffentlicht: (2026)
von: Wang, Kuanning, et al.
Veröffentlicht: (2026)
AffordDP: Generalizable Diffusion Policy with Transferable Affordance
von: Wu, Shijie, et al.
Veröffentlicht: (2024)
von: Wu, Shijie, et al.
Veröffentlicht: (2024)
You Only Estimate Once: Unified, One-stage, Real-Time Category-level Articulated Object 6D Pose Estimation for Robotic Grasping
von: Huang, Jingshun, et al.
Veröffentlicht: (2025)
von: Huang, Jingshun, et al.
Veröffentlicht: (2025)
TP-MDDN: Task-Preferenced Multi-Demand-Driven Navigation with Autonomous Decision-Making
von: Li, Shanshan, et al.
Veröffentlicht: (2025)
von: Li, Shanshan, et al.
Veröffentlicht: (2025)
OFlow: Injecting Object-Aware Temporal Flow Matching for Robust Robotic Manipulation
von: Wang, Kuanning, et al.
Veröffentlicht: (2026)
von: Wang, Kuanning, et al.
Veröffentlicht: (2026)
AffordDexGrasp: Open-set Language-guided Dexterous Grasp with Generalizable-Instructive Affordance
von: Wei, Yi-Lin, et al.
Veröffentlicht: (2025)
von: Wei, Yi-Lin, et al.
Veröffentlicht: (2025)
AffordTissue: Dense Affordance Prediction for Tool-Action Specific Tissue Interaction
von: Maksutova, Aiza, et al.
Veröffentlicht: (2026)
von: Maksutova, Aiza, et al.
Veröffentlicht: (2026)
CoA-VLA: Improving Vision-Language-Action Models via Visual-Textual Chain-of-Affordance
von: Li, Jinming, et al.
Veröffentlicht: (2024)
von: Li, Jinming, et al.
Veröffentlicht: (2024)
Universal Pose Pretraining for Generalizable Vision-Language-Action Policies
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
von: Lin, Haitao, et al.
Veröffentlicht: (2026)
SCOOP'D: Learning Mixed-Liquid-Solid Scooping via Sim2Real Generative Policy
von: Wang, Kuanning, et al.
Veröffentlicht: (2025)
von: Wang, Kuanning, et al.
Veröffentlicht: (2025)
ST4VLA: Spatially Guided Training for Vision-Language-Action Models
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
von: Ye, Jinhui, et al.
Veröffentlicht: (2026)
AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
von: Tang, Yingbo, et al.
Veröffentlicht: (2025)
von: Tang, Yingbo, et al.
Veröffentlicht: (2025)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
von: Wu, Xiaofei, et al.
Veröffentlicht: (2026)
von: Wu, Xiaofei, et al.
Veröffentlicht: (2026)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
von: Zhu, Xiaomeng, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaomeng, et al.
Veröffentlicht: (2025)
DropVLA: An Action-Level Backdoor Attack on Vision-Language-Action Models
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
von: Xu, Zonghuan, et al.
Veröffentlicht: (2025)
Spatial-Temporal Aware Visuomotor Diffusion Policy Learning
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
AffordSim: A Scalable Data Generator and Benchmark for Affordance-Aware Robotic Manipulation
von: Li, Mingyang, et al.
Veröffentlicht: (2026)
von: Li, Mingyang, et al.
Veröffentlicht: (2026)
RoboAfford++: A Generative AI-Enhanced Dataset for Multimodal Affordance Learning in Robotic Manipulation and Navigation
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025)
von: Hao, Xiaoshuai, et al.
Veröffentlicht: (2025)
Robotic Grasping and Placement Controlled by EEG-Based Hybrid Visual and Motor Imagery
von: Liu, Yichang, et al.
Veröffentlicht: (2026)
von: Liu, Yichang, et al.
Veröffentlicht: (2026)
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
von: Ding, Kairui, et al.
Veröffentlicht: (2024)
von: Ding, Kairui, et al.
Veröffentlicht: (2024)
SparseGrasp: Robotic Grasping via 3D Semantic Gaussian Splatting from Sparse Multi-View RGB Images
von: Yu, Junqiu, et al.
Veröffentlicht: (2024)
von: Yu, Junqiu, et al.
Veröffentlicht: (2024)
EqvAfford: SE(3) Equivariance for Point-Level Affordance Learning
von: Chen, Yue, et al.
Veröffentlicht: (2024)
von: Chen, Yue, et al.
Veröffentlicht: (2024)
CAP-Net: A Unified Network for 6D Pose and Size Estimation of Categorical Articulated Parts from a Single RGB-D Image
von: Huang, Jingshun, et al.
Veröffentlicht: (2025)
von: Huang, Jingshun, et al.
Veröffentlicht: (2025)
RynnVLA-002: A Unified Vision-Language-Action and World Model
von: Cen, Jun, et al.
Veröffentlicht: (2025)
von: Cen, Jun, et al.
Veröffentlicht: (2025)
VP-VLA: Visual Prompting as an Interface for Vision-Language-Action Models
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
von: Wang, Zixuan, et al.
Veröffentlicht: (2026)
Polaris: Open-ended Interactive Robotic Manipulation via Syn2Real Visual Grounding and Large Language Models
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
von: Wang, Tianyu, et al.
Veröffentlicht: (2024)
A Neural Representation Framework with LLM-Driven Spatial Reasoning for Open-Vocabulary 3D Visual Grounding
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025)
Joint-Aligned Latent Action: Towards Scalable VLA Pretraining in the Wild
von: Luo, Hao, et al.
Veröffentlicht: (2026)
von: Luo, Hao, et al.
Veröffentlicht: (2026)
AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Afford Correspondence
von: Zhang, Jiawei, et al.
Veröffentlicht: (2026)
von: Zhang, Jiawei, et al.
Veröffentlicht: (2026)
Sequential Multi-Object Grasping with One Dexterous Hand
von: He, Sicheng, et al.
Veröffentlicht: (2025)
von: He, Sicheng, et al.
Veröffentlicht: (2025)
Agile-VLA: Few-Shot Industrial Pose Rectification via Implicit Affordance Anchoring
von: Yan, Teng, et al.
Veröffentlicht: (2026)
von: Yan, Teng, et al.
Veröffentlicht: (2026)
HARP-VLA: Human-Robot Aligned Representation Learning for Vision-Language-Action Model
von: Zhu, Xiang, et al.
Veröffentlicht: (2026)
von: Zhu, Xiang, et al.
Veröffentlicht: (2026)
Diffusion-Based Imaginative Coordination for Bimanual Manipulation
von: Xu, Huilin, et al.
Veröffentlicht: (2025)
von: Xu, Huilin, et al.
Veröffentlicht: (2025)
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
von: Xiao, Lei, et al.
Veröffentlicht: (2025)
VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection
von: Sun, Haowen, et al.
Veröffentlicht: (2026)
von: Sun, Haowen, et al.
Veröffentlicht: (2026)
CoT-VLA: Visual Chain-of-Thought Reasoning for Vision-Language-Action Models
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
von: Zhao, Qingqing, et al.
Veröffentlicht: (2025)
RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models
von: Zhang, Hongyin, et al.
Veröffentlicht: (2025)
von: Zhang, Hongyin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TriVLA: A Triple-System-Based Unified Vision-Language-Action Model with Episodic World Modeling for General Robot Control
von: Liu, Zhenyang, et al.
Veröffentlicht: (2025) -
AffordVLA: Injecting Affordance Representations into Vision-Language-Action Models via Implicit Feature Alignment
von: Kong, Weijie, et al.
Veröffentlicht: (2026) -
ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
von: Liu, Zhenyang, et al.
Veröffentlicht: (2026) -
OCRA: Object-Centric Learning with 3D and Tactile Priors for Human-to-Robot Action Transfer
von: Wang, Kuanning, et al.
Veröffentlicht: (2026) -
AffordDP: Generalizable Diffusion Policy with Transferable Affordance
von: Wu, Shijie, et al.
Veröffentlicht: (2024)