Affordance Agent Harness: Verification-Gated Skill Orchestration
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Haojian, Shi, Jiahao, Li, Yinchuan, Chen, Yingcong |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
por: Chu, Hengshuo, et al.
Publicado: (2025)
por: Chu, Hengshuo, et al.
Publicado: (2025)
Find, Fix, Reason: Context Repair for Video Reasoning
por: Huang, Haojian, et al.
Publicado: (2026)
por: Huang, Haojian, et al.
Publicado: (2026)
Panoramic Affordance Prediction
por: Zhang, Zixin, et al.
Publicado: (2026)
por: Zhang, Zixin, et al.
Publicado: (2026)
AffordanceLLM: Grounding Affordance from Vision Language Models
por: Qian, Shengyi, et al.
Publicado: (2024)
por: Qian, Shengyi, et al.
Publicado: (2024)
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
por: Wang, Hanqing, et al.
Publicado: (2025)
por: Wang, Hanqing, et al.
Publicado: (2025)
A4-Agent: An Agentic Framework for Zero-Shot Affordance Reasoning
por: Zhang, Zixin, et al.
Publicado: (2025)
por: Zhang, Zixin, et al.
Publicado: (2025)
AnySkill: Learning Open-Vocabulary Physical Skill for Interactive Agents
por: Cui, Jieming, et al.
Publicado: (2024)
por: Cui, Jieming, et al.
Publicado: (2024)
AffordGrasp: Cross-Modal Diffusion for Affordance-Aware Grasp Synthesis
por: Wu, Xiaofei, et al.
Publicado: (2026)
por: Wu, Xiaofei, et al.
Publicado: (2026)
Egocentric Instruction-oriented Affordance Prediction via Large Multimodal Model
por: Ji, Bokai, et al.
Publicado: (2025)
por: Ji, Bokai, et al.
Publicado: (2025)
AffordanceGrasp-R1:Leveraging Reasoning-Based Affordance Segmentation with Reinforcement Learning for Robotic Grasping
por: Zhou, Dingyi, et al.
Publicado: (2026)
por: Zhou, Dingyi, et al.
Publicado: (2026)
TRACER: Texture-Robust Affordance Chain-of-Thought for Deformable-Object Refinement
por: Jia, Wanjun, et al.
Publicado: (2026)
por: Jia, Wanjun, et al.
Publicado: (2026)
PAVLM: Advancing Point Cloud based Affordance Understanding Via Vision-Language Model
por: Liu, Shang-Ching, et al.
Publicado: (2024)
por: Liu, Shang-Ching, et al.
Publicado: (2024)
Visual Affordance Prediction: Survey and Reproducibility
por: Apicella, Tommaso, et al.
Publicado: (2025)
por: Apicella, Tommaso, et al.
Publicado: (2025)
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
por: Ding, Kairui, et al.
Publicado: (2024)
por: Ding, Kairui, et al.
Publicado: (2024)
CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects
por: Li, Jingliang, et al.
Publicado: (2026)
por: Li, Jingliang, et al.
Publicado: (2026)
Afford-X: Generalizable and Slim Affordance Reasoning for Task-oriented Manipulation
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
por: Zhu, Xiaomeng, et al.
Publicado: (2025)
HRP: Human Affordances for Robotic Pre-Training
por: Srirama, Mohan Kumar, et al.
Publicado: (2024)
por: Srirama, Mohan Kumar, et al.
Publicado: (2024)
Resource-Efficient Affordance Grounding with Complementary Depth and Semantic Prompts
por: Huang, Yizhou, et al.
Publicado: (2025)
por: Huang, Yizhou, et al.
Publicado: (2025)
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation
por: Li, Gen, et al.
Publicado: (2024)
por: Li, Gen, et al.
Publicado: (2024)
Interpretable Affordance Detection on 3D Point Clouds with Probabilistic Prototypes
por: Li, Maximilian Xiling, et al.
Publicado: (2025)
por: Li, Maximilian Xiling, et al.
Publicado: (2025)
Ag2Manip: Learning Novel Manipulation Skills with Agent-Agnostic Visual and Action Representations
por: Li, Puhao, et al.
Publicado: (2024)
por: Li, Puhao, et al.
Publicado: (2024)
VoxAfford: Multi-Scale Voxel-Token Fusion for Open-Vocabulary 3D Affordance Detection
por: Sun, Haowen, et al.
Publicado: (2026)
por: Sun, Haowen, et al.
Publicado: (2026)
RoboPCA: Pose-centered Affordance Learning from Human Demonstrations for Robot Manipulation
por: Xiao, Zhanqi, et al.
Publicado: (2026)
por: Xiao, Zhanqi, et al.
Publicado: (2026)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
por: Korekata, Ryosuke, et al.
Publicado: (2025)
por: Korekata, Ryosuke, et al.
Publicado: (2025)
ManipGPT: Is Affordance Segmentation by Large Vision Models Enough for Articulated Object Manipulation?
por: Kim, Taewhan, et al.
Publicado: (2024)
por: Kim, Taewhan, et al.
Publicado: (2024)
UAD: Unsupervised Affordance Distillation for Generalization in Robotic Manipulation
por: Tang, Yihe, et al.
Publicado: (2025)
por: Tang, Yihe, et al.
Publicado: (2025)
DAP: Diffusion-based Affordance Prediction for Multi-modality Storage
por: Chang, Haonan, et al.
Publicado: (2024)
por: Chang, Haonan, et al.
Publicado: (2024)
Simultaneous Localization and Affordance Prediction of Tasks from Egocentric Video
por: Chavis, Zachary, et al.
Publicado: (2024)
por: Chavis, Zachary, et al.
Publicado: (2024)
RAGNet: Large-scale Reasoning-based Affordance Segmentation Benchmark towards General Grasping
por: Wu, Dongming, et al.
Publicado: (2025)
por: Wu, Dongming, et al.
Publicado: (2025)
NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation
por: Xu, Ran, et al.
Publicado: (2024)
por: Xu, Ran, et al.
Publicado: (2024)
Agentic Scene Policies: Unifying Space, Semantics, and Affordances for Robot Action
por: Morin, Sacha, et al.
Publicado: (2025)
por: Morin, Sacha, et al.
Publicado: (2025)
GLOVER: Generalizable Open-Vocabulary Affordance Reasoning for Task-Oriented Grasping
por: Ma, Teli, et al.
Publicado: (2024)
por: Ma, Teli, et al.
Publicado: (2024)
PhysToolBench: Benchmarking Physical Tool Understanding for MLLMs
por: Zhang, Zixin, et al.
Publicado: (2025)
por: Zhang, Zixin, et al.
Publicado: (2025)
RAM: Retrieval-Based Affordance Transfer for Generalizable Zero-Shot Robotic Manipulation
por: Kuang, Yuxuan, et al.
Publicado: (2024)
por: Kuang, Yuxuan, et al.
Publicado: (2024)
GLOVER++: Unleashing the Potential of Affordance Learning from Human Behaviors for Robotic Manipulation
por: Ma, Teli, et al.
Publicado: (2025)
por: Ma, Teli, et al.
Publicado: (2025)
RING#: PR-by-PE Global Localization with Roto-translation Equivariant Gram Learning
por: Lu, Sha, et al.
Publicado: (2024)
por: Lu, Sha, et al.
Publicado: (2024)
PanoAffordanceNet: Towards Holistic Affordance Grounding in 360° Indoor Environments
por: Zhu, Guoliang, et al.
Publicado: (2026)
por: Zhu, Guoliang, et al.
Publicado: (2026)
ModSkill: Physical Character Skill Modularization
por: Huang, Yiming, et al.
Publicado: (2025)
por: Huang, Yiming, et al.
Publicado: (2025)
Robo-ABC: Affordance Generalization Beyond Categories via Semantic Correspondence for Robot Manipulation
por: Ju, Yuanchen, et al.
Publicado: (2024)
por: Ju, Yuanchen, et al.
Publicado: (2024)
TextInPlace: Indoor Visual Place Recognition in Repetitive Structures with Scene Text Spotting and Verification
por: Tao, Huaqi, et al.
Publicado: (2025)
por: Tao, Huaqi, et al.
Publicado: (2025)
Ejemplares similares
-
3D-AffordanceLLM: Harnessing Large Language Models for Open-Vocabulary Affordance Detection in 3D Worlds
por: Chu, Hengshuo, et al.
Publicado: (2025) -
Find, Fix, Reason: Context Repair for Video Reasoning
por: Huang, Haojian, et al.
Publicado: (2026) -
Panoramic Affordance Prediction
por: Zhang, Zixin, et al.
Publicado: (2026) -
AffordanceLLM: Grounding Affordance from Vision Language Models
por: Qian, Shengyi, et al.
Publicado: (2024) -
Affordance-R1: Reinforcement Learning for Generalizable Affordance Reasoning in Multimodal Large Language Model
por: Wang, Hanqing, et al.
Publicado: (2025)