DAP: Diffusion-based Affordance Prediction for Multi-modality Storage
Fuente:
arXiv
Saved in:
| Main Authors: | Chang, Haonan, Boyalakuntla, Kowndinya, Liu, Yuhan, Zhang, Xinyu, Schramm, Liam, Boularias, Abdeslam |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Motion Blender Gaussian Splatting for Dynamic Scene Reconstruction
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
KARL: Kalman-Filter Assisted Reinforcement Learner for Dynamic Object Tracking and Grasping
by: Boyalakuntla, Kowndinya, et al.
Published: (2025)
by: Boyalakuntla, Kowndinya, et al.
Published: (2025)
Bellman Diffusion Models
by: Schramm, Liam, et al.
Published: (2024)
by: Schramm, Liam, et al.
Published: (2024)
Autoregressive Action Sequence Learning for Robotic Manipulation
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
Detect Everything with Few Examples
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
Provably Efficient Long-Horizon Exploration in Monte Carlo Tree Search through State Occupancy Regularization
by: Schramm, Liam, et al.
Published: (2024)
by: Schramm, Liam, et al.
Published: (2024)
Scaling Manipulation Learning with Visual Kinematic Chain Prediction
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
Learning Visual Feature-Based World Models via Residual Latent Action
by: Zhang, Xinyu, et al.
Published: (2026)
by: Zhang, Xinyu, et al.
Published: (2026)
Failure Forecasting Boosts Robustness of Sim2Real Rhythmic Insertion Policies
by: Liu, Yuhan, et al.
Published: (2025)
by: Liu, Yuhan, et al.
Published: (2025)
LGMCTS: Language-Guided Monte-Carlo Tree Search for Executable Semantic Object Rearrangement
by: Chang, Haonan, et al.
Published: (2023)
by: Chang, Haonan, et al.
Published: (2023)
Panoramic Affordance Prediction
by: Zhang, Zixin, et al.
Published: (2026)
by: Zhang, Zixin, et al.
Published: (2026)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
by: Wu, Xiaofei, et al.
Published: (2026)
by: Wu, Xiaofei, et al.
Published: (2026)
BEACON: Language-Conditioned Navigation Affordance Prediction under Occlusion
by: Gao, Xinyu, et al.
Published: (2026)
by: Gao, Xinyu, et al.
Published: (2026)
Visual Affordance Prediction: Survey and Reproducibility
by: Apicella, Tommaso, et al.
Published: (2025)
by: Apicella, Tommaso, et al.
Published: (2025)
Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
by: Morin, Sacha, et al.
Published: (2025)
by: Morin, Sacha, et al.
Published: (2025)
MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning
by: Chen, Guangli, et al.
Published: (2026)
by: Chen, Guangli, et al.
Published: (2026)
Multi-modal Motion Prediction using Temporal Ensembling with Learning-based Aggregation
by: Hong, Kai-Yin, et al.
Published: (2024)
by: Hong, Kai-Yin, et al.
Published: (2024)
Simultaneous Localization and Affordance Prediction of Tasks from Egocentric Video
by: Chavis, Zachary, et al.
Published: (2024)
by: Chavis, Zachary, et al.
Published: (2024)
AffordanceLLM: Grounding Affordance from Vision Language Models
by: Qian, Shengyi, et al.
Published: (2024)
by: Qian, Shengyi, et al.
Published: (2024)
Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model
by: Ji, Bokai, et al.
Published: (2025)
by: Ji, Bokai, et al.
Published: (2025)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
by: Wang, Hanqing, et al.
Published: (2025)
by: Wang, Hanqing, et al.
Published: (2025)
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
by: Zhou, Dingyi, et al.
Published: (2026)
by: Zhou, Dingyi, et al.
Published: (2026)
PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model
by: Liu, Shang-Ching, et al.
Published: (2024)
by: Liu, Shang-Ching, et al.
Published: (2024)
Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation
by: Xie, Zongwu, et al.
Published: (2025)
by: Xie, Zongwu, et al.
Published: (2025)
VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection
by: Sun, Haowen, et al.
Published: (2026)
by: Sun, Haowen, et al.
Published: (2026)
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
by: Wu, Dongming, et al.
Published: (2025)
by: Wu, Dongming, et al.
Published: (2025)
MM-Gaussian: 3D Gaussian-based Multi-modal Fusion for Localization and Reconstruction in Unbounded Scenes
by: Wu, Chenyang, et al.
Published: (2024)
by: Wu, Chenyang, et al.
Published: (2024)
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
by: Chu, Hengshuo, et al.
Published: (2025)
by: Chu, Hengshuo, et al.
Published: (2025)
HRP: Human Affordances for Robotic Pre-Training
by: Srirama, Mohan Kumar, et al.
Published: (2024)
by: Srirama, Mohan Kumar, et al.
Published: (2024)
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
by: Zhang, Zixin, et al.
Published: (2025)
by: Zhang, Zixin, et al.
Published: (2025)
MoMa-Kitchen: A 100K+ Benchmark for Affordance-Grounded Last-Mile Navigation in Mobile Manipulation
by: Zhang, Pingrui, et al.
Published: (2025)
by: Zhang, Pingrui, et al.
Published: (2025)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
by: Korekata, Ryosuke, et al.
Published: (2025)
by: Korekata, Ryosuke, et al.
Published: (2025)
Affordance Agent Harness: Verification-Gated Skill Orchestration
by: Huang, Haojian, et al.
Published: (2026)
by: Huang, Haojian, et al.
Published: (2026)
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
by: Ju, Yuanchen, et al.
Published: (2024)
by: Ju, Yuanchen, et al.
Published: (2024)
SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
by: Wu, Shengkai, et al.
Published: (2025)
by: Wu, Shengkai, et al.
Published: (2025)
TRACER: Texture-Robust Affordance Chain-of-Thought for Deformable-Object Refinement
by: Jia, Wanjun, et al.
Published: (2026)
by: Jia, Wanjun, et al.
Published: (2026)
One-Shot Imitation Learning with Invariance Matching for Robotic Manipulation
by: Zhang, Xinyu, et al.
Published: (2024)
by: Zhang, Xinyu, et al.
Published: (2024)
CLIP-Loc: Multi-modal Landmark Association for Global Localization in Object-based Maps
by: Matsuzaki, Shigemichi, et al.
Published: (2024)
by: Matsuzaki, Shigemichi, et al.
Published: (2024)
Distilling Multi-modal Large Language Models for Autonomous Driving
by: Hegde, Deepti, et al.
Published: (2025)
by: Hegde, Deepti, et al.
Published: (2025)
Similar Items
-
Motion Blender Gaussian Splatting for Dynamic Scene Reconstruction
by: Zhang, Xinyu, et al.
Published: (2025) -
KARL: Kalman-Filter Assisted Reinforcement Learner for Dynamic Object Tracking and Grasping
by: Boyalakuntla, Kowndinya, et al.
Published: (2025) -
Bellman Diffusion Models
by: Schramm, Liam, et al.
Published: (2024) -
Autoregressive Action Sequence Learning for Robotic Manipulation
by: Zhang, Xinyu, et al.
Published: (2024) -
Detect Everything with Few Examples
by: Zhang, Xinyu, et al.
Published: (2023)